Lucene search
+L

1 matches found

Packet Storm News
Packet Storm News
added 2025/05/26 12:0 a.m.10 views

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks

Large language models LLMs have achieved remarkable capabilities but remain vulnerable to adversarial prompts known as jailbreaks, which can bypass safety alignment and elicit harmful outputs. Despite growing efforts in LLM safety research, existing evaluations are often fragmented, focused on...

7.3AI score
SaveExploits0
Rows per page
Query Builder