Lucene search
+L

2 matches found

Kitploit
Kitploit
•added 2026/09/27 11:45 p.m.•14 views

llm-censorship-steering

LLM Censorship Steering This repository contains code implementation for Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control by Hannah Cyberey and David Evans. We introduce a method that finds "steering vectors" from LLM internals for detecting and controlling the...

6.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/04/17 12:00 a.m.•47 views

Surgical Repair of Insecure Code Generation in LLMs

Large language models write production code, and yet they routinely introduce well-known vulnerabilities. We show that this is not a knowledge deficit: the same models that generate insecure code, correctly identify and explain the vulnerability when asked directly, this is a gap we call the...

5.8AI score
SaveExploits0
Rows per page
Query Builder