Lucene search
+L

30 matches found

Packet Storm News
Packet Storm News
•added 2026/10/06 12:00 a.m.•7 views

Secure Speculative Decoding for Large Language Models

Speculative decoding accelerates inference for a large language model LLM, referred to as the target model, by first using a smaller model, referred to as the draft model, to generate candidate tokens and then verifying them with the target model for acceptance or rejection. Prior studies primari...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/10/05 12:00 a.m.•6 views

Safeguarding LLMs Via Model-Agnostic Latent Safety Signals from Dark Knowledge

LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models' hidden states. However, existing decoding-stage defenses suffer from two limitations. First,...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/07/19 12:00 a.m.•37 views

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions

Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate ASR, yet an attack that breaks a model is not necessarily useful for improving its safety. We propose a defender-centric view of jailbreak evaluation, where attacks are evaluat...

5.4AI score
SaveExploits0
OSV
OSV
•added 2026/07/07 4:02 p.m.•24 views

PYSEC-2026-1624 Mesop Class Pollution vulnerability leads to DoS and Jailbreak attacks

From @jackfromeast and @superboy-zjc: We have identified a class pollution vulnerability in Mesop = 0.14.0 application that allows attackers to overwrite global variables and class attributes in certain Mesop modules during runtime. This vulnerability could directly lead to a denial of service Do...

8.1CVSS6AI score0.00704EPSS
SaveExploits0References6
Packet Storm News
Packet Storm News
•added 2026/07/03 12:00 a.m.•18 views

Overloading Large Vision-Language Models for Jailbreaking

Large Vision-Language Models LVLMs exhibit remarkable vision-language capabilities and are increasingly deployed in real-world applications such as personal assistants, document analysis systems, and embodied agents. However, their dual-modal attack surfaces make them vulnerable to jailbreak...

6AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/06/16 12:00 a.m.•16 views

A Red-Team Study of Anthropic Fable 5 and Opus 4.8 Models

We evaluate the adversarial robustness of two frontier large language models LLMs developed by Anthropic, Fable 5 and Opus 4.8, against four families of automated jailbreak attack across 7 826 harmful intents spanning a ten-category harm taxonomy. Using the HackAgent red-teaming framework, hundre...

5.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/02/24 12:00 a.m.•19 views

Analysis of LLMs against Prompt Injection and Jailbreak Attacks

Large Language Models LLMs are widely deployed in real-world systems. Given their broader applicability, prompt engineering has become an efficient tool for resource-scarce organizations to adopt LLMs for their own purposes. At the same time, LLMs are vulnerable to prompt-based attacks. Thus,...

6AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/02/06 12:00 a.m.•15 views

ShallowJail: Steering Jailbreaks against Large Language Models

Large Language ModelsLLMs have been successful in numerous fields. Alignment has usually been applied to prevent them from harmful purposes. However, aligned LLMs remain vulnerable to jailbreak attacks that deliberately mislead them into producing harmful outputs. Existing jailbreaks are either...

5.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/01/07 12:00 a.m.•39 views

Jailbreaking LLMs and VLMs: Mechanisms, Evaluation, and Unified Defense

This paper provides a systematic survey of jailbreak attacks and defenses on Large Language Models LLMs and Vision-Language Models VLMs, emphasizing that jailbreak vulnerabilities stem from structural factors such as incomplete training data, linguistic ambiguity, and generative uncertainty. It...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/12/23 12:00 a.m.•42 views

Odysseus: Jailbreaking Commercial Multimodal LLM-Integrated Systems Via Dual Steganography

By integrating language understanding with perceptual modalities such as images, multimodal large language models MLLMs constitute a critical substrate for modern AI systems, particularly intelligent agents operating in open and interactive environments. However, their increasing accessibility al...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/11/14 12:00 a.m.•17 views

NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks

Jailbreak attacks designed to bypass safety mechanisms pose a serious threat by prompting LLMs to generate harmful or inappropriate content, despite alignment with ethical guidelines. Crafting universal filtering rules remains difficult due to their inherent dependence on specific contexts. To...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/10/21 12:00 a.m.•24 views

HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models

Large Language Models LLMs remain vulnerable to multi-turn jailbreak attacks. We introduce HarmNet, a modular framework comprising ThoughtNet, a hierarchical semantic network; a feedback-driven Simulator for iterative query refinement; and a Network Traverser for real-time adaptive attack...

7.1AI score
SaveExploits0
EUVD
EUVD
•added 2025/10/03 8:07 p.m.•38 views

EUVD-2025-14804

Malicious code in bioql PyPI...

8.1CVSS6.4AI score0.00704EPSS
SaveExploits0References3
Packet Storm News
Packet Storm News
•added 2025/10/01 12:00 a.m.•56 views

Breaking the Code: Security Assessment of AI Code Agents through Systematic Jailbreaking Attacks

Code-capable large language model LLM agents are increasingly embedded into software engineering workflows where they can read, write, and execute code, raising the stakes of safety-bypass "jailbreak" attacks beyond text-only settings. Prior evaluations emphasize refusal or harmful-text detection...

7.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/09/18 12:00 a.m.•20 views

Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism Via Probabilistically Ablating Refusal Direction

Jailbreak attacks pose persistent threats to large language models LLMs. Current safety alignment methods have attempted to address these issues, but they experience two significant limitations: insufficient safety alignment depth and unrobust internal defense mechanisms. These limitations make...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/09/08 12:00 a.m.•17 views

Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?

Jailbreak attacks on Large Language Models LLMs have demonstrated various successful methods whereby attackers manipulate models into generating harmful responses that they are designed to avoid. Among these, Greedy Coordinate Gradient GCG has emerged as a general and effective approach that...

7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/09/04 12:00 a.m.•16 views

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

In deployment and application, large language models LLMs typically undergo safety alignment to prevent illegal and unethical outputs. However, the continuous advancement of jailbreak attack techniques, designed to bypass safety mechanisms with adversarial prompts, has placed increasing pressure ...

7.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/08/16 12:00 a.m.•15 views

Mitigating Jailbreaks with Intent-Aware LLMs

Despite extensive safety-tuning, large language models LLMs remain vulnerable to jailbreak attacks via adversarially crafted instructions, reflecting a persistent trade-off between safety and task performance. In this work, we propose Intent-FT, a simple and lightweight fine-tuning approach that...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/07/06 12:00 a.m.•19 views

Attention Slipping: a Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

As large language models LLMs become more integral to society and technology, ensuring their safety becomes essential. Jailbreak attacks exploit vulnerabilities to bypass safety guardrails, posing a significant threat. However, the mechanisms enabling these attacks are not well understood. In thi...

7.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/23 12:00 a.m.•15 views

Security Assessment of DeepSeek and GPT Series Models against Jailbreak Attacks

The widespread deployment of large language models LLMs has raised critical concerns over their vulnerability to jailbreak attacks, i.e., adversarial prompts that bypass alignment mechanisms and elicit harmful or policy-violating outputs. While proprietary models like GPT-4 have undergone extensi...

7.3AI score
SaveExploits0
Rows per page
Query Builder