2 matches found
llm-censorship-steering
LLM Censorship Steering This repository contains code implementation for Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control by Hannah Cyberey and David Evans. We introduce a method that finds "steering vectors" from LLM internals for detecting and controlling the...
5.9AI score
SaveExploits0
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
Large language models LLMs have transformed the way we access information. These models are often tuned to refuse to comply with requests that are considered harmful and to produce responses that better align with the preferences of those who control the models. To understand how this "censorship...
7.2AI score
SaveExploits0
20