Lucene search
+L

379 matches found

Packet Storm News
Packet Storm News
added 2025/10/01 12:0 a.m.12 views

Breaking the Code: Security Assessment of AI Code Agents through Systematic Jailbreaking Attacks

Code-capable large language model LLM agents are increasingly embedded into software engineering workflows where they can read, write, and execute code, raising the stakes of safety-bypass "jailbreak" attacks beyond text-only settings. Prior evaluations emphasize refusal or harmful-text detection...

7.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/09/28 12:0 a.m.5 views

Quant Fever, Reasoning Blackholes, Schrodinger'S Compliance, and More: Probing GPT-OSS-20B

OpenAI's GPT-OSS family provides open-weight language models with explicit chain-of-thought CoT reasoning and a Harmony prompt format. We summarize an extensive security evaluation of GPT-OSS-20B that probes the model's behavior under different adversarial conditions. Using the Jailbreak Oracle J...

6.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/09/24 12:0 a.m.11 views

Bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs

With the rapid advancement of large language models LLMs, their robustness against adversarial manipulations, particularly jailbreak backdoor attacks, has become critically important. Existing approaches to embedding jailbreak triggers--such as supervised fine-tuning SFT, model editing, and...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/09/20 12:0 a.m.7 views

DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems

Intelligent software systems powered by Large Language Models LLMs are increasingly deployed in critical sectors, raising concerns about their safety during runtime. Through an industry-academic collaboration when deploying an LLM-powered virtual customer assistant, a critical software engineerin...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/09/18 12:0 a.m.7 views

Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism Via Probabilistically Ablating Refusal Direction

Jailbreak attacks pose persistent threats to large language models LLMs. Current safety alignment methods have attempted to address these issues, but they experience two significant limitations: insufficient safety alignment depth and unrobust internal defense mechanisms. These limitations make...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/09/17 12:0 a.m.3 views

LLM Jailbreak Detection for (Almost) Free!

Large language models LLMs enhance security through alignment when widely used, but remain susceptible to jailbreak attacks capable of producing inappropriate content. Jailbreak detection methods show promise in mitigating jailbreak attacks through the assistance of other models or multiple model...

6.8AI score
SaveExploits0
Gitee
Gitee
added 2025/09/14 4:34 p.m.122 views

Exploit for CVE-2016-4655

This is a PoC exploit for iOS 9.3.5, targeting CVE-2016-4655 and CVE-2016-4656. The exploit aims to gain root access over the device by exploiting kernel vulnerabilities. The supported devices are listed in offsetfinder.h. The exploit is based on the original disclosure by Lookout and the OS X...

9.3CVSS7.1AI score0.63579EPSS
SaveExploits13
Packet Storm News
Packet Storm News
added 2025/09/08 12:0 a.m.7 views

Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?

Jailbreak attacks on Large Language Models LLMs have demonstrated various successful methods whereby attackers manipulate models into generating harmful responses that they are designed to avoid. Among these, Greedy Coordinate Gradient GCG has emerged as a general and effective approach that...

7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/09/04 12:0 a.m.7 views

Between a Rock and a Hard Place: Exploiting Ethical Reasoning to Jailbreak LLMs

Large language models LLMs have undergone safety alignment efforts to mitigate harmful outputs. However, as LLMs become more sophisticated in reasoning, their intelligence may introduce new security risks. While traditional jailbreak attacks relied on singlestep attacks, multi-turn jailbreak...

7.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/09/04 12:0 a.m.8 views

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

In deployment and application, large language models LLMs typically undergo safety alignment to prevent illegal and unethical outputs. However, the continuous advancement of jailbreak attack techniques, designed to bypass safety mechanisms with adversarial prompts, has placed increasing pressure ...

7.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/08/22 12:0 a.m.4 views

Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models

Large Language Models LLMs remain vulnerable to jailbreak attacks, which attempt to elicit harmful responses from LLMs. The evolving nature and diversity of these attacks pose many challenges for defense systems, including 1 adaptation to counter emerging attack strategies without costly...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/08/16 12:0 a.m.6 views

Mitigating Jailbreaks with Intent-Aware LLMs

Despite extensive safety-tuning, large language models LLMs remain vulnerable to jailbreak attacks via adversarially crafted instructions, reflecting a persistent trade-off between safety and task performance. In this work, we propose Intent-FT, a simple and lightweight fine-tuning approach that...

7.2AI score
SaveExploits0
OSSF Malicious Packages
OSSF Malicious Packages
added 2025/08/14 6:52 p.m.2 views

Malicious code in jailbreak-ios (npm)

The package jailbreak-ios was found to contain malicious code...

7AI score
SaveExploits0
OSSF Malicious Packages
OSSF Malicious Packages
added 2025/08/14 6:52 p.m.3 views

Malicious code in is-jailbreak (npm)

The package is-jailbreak was found to contain malicious code...

7AI score
SaveExploits0
OSV
OSV
added 2025/08/14 6:52 p.m.3 views

MAL-2025-23394 Malicious code in is-jailbreak (npm)

The package is-jailbreak was found to contain malicious code...

7.2AI score
SaveExploits0
OSV
OSV
added 2025/08/14 6:52 p.m.2 views

MAL-2025-23583 Malicious code in jailbreak-ios (npm)

The package jailbreak-ios was found to contain malicious code...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/08/14 12:0 a.m.2 views

Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts

Evaluating jailbreak attacks is challenging when prompts are not overtly harmful or fail to induce harmful outputs. Unfortunately, many existing red-teaming datasets contain such unsuitable prompts. To evaluate attacks accurately, these datasets need to be assessed and cleaned for maliciousness...

6.9AI score
SaveExploits0
The Hacker News
The Hacker News
added 2025/08/09 3:6 p.m.16 views

Researchers Uncover GPT-5 Jailbreak and Zero-Click AI Agent Attacks Exposing Cloud and IoT Systems

Cybersecurity researchers have uncovered a jailbreak technique to bypass ethical guardrails erected by OpenAI in its latest large language model LLM GPT-5 and produce illicit instructions. Generative artificial intelligence AI security platform NeuralTrust said it combined a known technique calle...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/08/08 12:0 a.m.10 views

Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs

Large language models LLMs demonstrate impressive capabilities in various language tasks but are susceptible to jailbreak attacks that circumvent their safety alignments. This paper introduces Latent Fusion Jailbreak LFJ, a representation-based attack that interpolates hidden states from harmful...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/08/08 12:0 a.m.4 views

Beyond Uniform Criteria: Scenario-Adaptive Multi-Dimensional Jailbreak Evaluation

Precise jailbreak evaluation is vital for LLM red teaming and jailbreak research. Current approaches employ binary classification e.g., string matching, toxic text classifiers, LLM-driven methods, yielding only "yes/no" labels without quantifying harm intensity. Existing multi-dimensional...

7AI score
SaveExploits0
Rows per page
Query Builder