Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2026/05/04 12:0 a.m.8 views

Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses

Defending large language models LLMs against jailbreak attacks, such as Greedy Coordinate Gradient GCG, remains a challenge, particularly under adaptive threat models where an attacker directly targets the defense mechanism. JBShield, a recent jailbreak defense with a 0% attack success rate in so...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/11/24 12:0 a.m.3 views

Defending Large Language Models against Jailbreak Exploits with Responsible AI Considerations

Large Language Models LLMs remain susceptible to jailbreak exploits that bypass safety filters and induce harmful or unethical behavior. This work presents a systematic taxonomy of existing jailbreak defenses across prompt-level, model-level, and training-time interventions, followed by three...

7.3AI score
SaveExploits0
Rows per page
Query Builder