Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
•added 2026/09/03 12:00 a.m.•8 views

AlcaTRAz - Anchored Tree-Rule Defense against Jailbreaks

Large language models LLMs are vulnerable to jailbreak attacks that bypass safety alignment through carefully crafted prompts. Many existing defenses require access to model weights or internals, making them difficult to apply to black-box deployments. We propose AlcaTRAz Anchored Tree-Rule defen...

5.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/05/11 12:00 a.m.•28 views

Guaranteed Jailbreaking Defense Via Disrupt-And-Rectify Smoothing

This paper proposes a guaranteed defense method for large language models LLMs to safeguard against jailbreaking attacks. Drawing inspiration from the denoised-smoothing approach in the adversarial defense domain, we propose a novel smoothing-based defense method, termed Disrupt-and-Rectify...

5.8AI score
SaveExploits0
Rows per page
Query Builder