1 matches found
Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
Aligned language models refuse harmful requests, but a one-line prefill "Sure, here is" strips the refusal. We ask where and how it fails. The harm representation stays intact: on the prompts the attack flips to compliance, a linear probe reads harm as high as on the refused ones 0.91-0.98, while...
5.4AI score
SaveExploits0
20