Lucene search
+L

3 matches found

Packet Storm News
Packet Storm News
added 2026/05/26 12:0 a.m.18 views

Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security

Large Language Models LLMs are increasingly vulnerable to adversarial prompts that exploit semantic ambiguities to bypass safety mechanisms, resulting in harmful or inappropriate outputs. Such attacks, including jailbreaking and prompt injection, pose significant risks to the integrity and...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/11 12:0 a.m.12 views

Re-Triggering Safeguards within LLMs for Jailbreak Detection

This paper proposes a jailbreaking prompt detection method for large language models LLMs to defend against jailbreak attacks. Although recent LLMs are equipped with built-in safeguards, it remains possible to craft jailbreaking prompts that bypass them. We argue that such jailbreaking prompts ar...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/16 12:0 a.m.6 views

LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs

Efficient red-teaming method to uncover vulnerabilities in Large Language Models LLMs is crucial. While recent attacks often use LLMs as optimizers, the discrete language space make gradient-based methods struggle. We introduce LARGO Latent Adversarial Reflection through Gradient Optimization, a...

7AI score
SaveExploits0
Rows per page
Query Builder