Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2026/03/15 12:0 a.m.4 views

Activation Surgery: Jailbreaking White-Box LLMs without Touching the Prompt

Most jailbreak techniques for Large Language Models LLMs primarily rely on prompt modifications, including paraphrasing, obfuscation, or conversational strategies. Meanwhile, abliteration techniques also known as targeted ablations of internal components have been used to study and explain LLM...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/04/26 12:0 a.m.7 views

Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control

Large language models LLMs have transformed the way we access information. These models are often tuned to refuse to comply with requests that are considered harmful and to produce responses that better align with the preferences of those who control the models. To understand how this "censorship...

7.2AI score
SaveExploits0
Rows per page
Query Builder