Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2025/09/17 12:0 a.m.4 views

LLM Jailbreak Detection for (Almost) Free!

Large language models LLMs enhance security through alignment when widely used, but remain susceptible to jailbreak attacks capable of producing inappropriate content. Jailbreak detection methods show promise in mitigating jailbreak attacks through the assistance of other models or multiple model...

6.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/07/06 12:0 a.m.6 views

Attention Slipping: a Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

As large language models LLMs become more integral to society and technology, ensuring their safety becomes essential. Jailbreak attacks exploit vulnerabilities to bypass safety guardrails, posing a significant threat. However, the mechanisms enabling these attacks are not well understood. In thi...

7.4AI score
SaveExploits0
Rows per page
Query Builder