1 matches found
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-Offs
Jailbreak defenses are essential for protecting large language models LLMs, but they can also introduce secondary costs that weaken model utility. We present a systematic study of these defense trade-offs along three dimensions: performance impact, over-refusal on benign inputs, and inference cos...
5.9AI score
SaveExploits0
20