Lucene search
+L

1 matches found

Packet Storm News
Packet Storm News
•added 2026/07/14 12:00 a.m.•12 views

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak

Aligned language models refuse harmful requests, but a one-line prefill "Sure, here is" strips the refusal. We ask where and how it fails. The harm representation stays intact: on the prompts the attack flips to compliance, a linear probe reads harm as high as on the refused ones 0.91-0.98, while...

5.4AI score
SaveExploits0
Rows per page
Query Builder