2 matches found
llm-censorship-steering
LLM Censorship Steering This repository contains code implementation for Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control by Hannah Cyberey and David Evans. We introduce a method that finds "steering vectors" from LLM internals for detecting and controlling the...
6.2AI score
SaveExploits0
Surgical Repair of Insecure Code Generation in LLMs
Large language models write production code, and yet they routinely introduce well-known vulnerabilities. We show that this is not a knowledge deficit: the same models that generate insecure code, correctly identify and explain the vulnerability when asked directly, this is a gap we call the...
5.8AI score
SaveExploits0
20