Lucene search
+L

379 matches found

Packet Storm News
Packet Storm News
added 2025/08/04 12:0 a.m.7 views

Large Reasoning Models Are Autonomous Jailbreak Agents

Jailbreaking -- bypassing built-in safety mechanisms in AI models -- has traditionally required complex technical procedures or specialized human expertise. In this study, we show that the persuasive capabilities of large reasoning models LRMs simplify and scale jailbreaking, converting it into a...

7.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/07/29 12:0 a.m.3 views

PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking

The increasing sophistication of large vision-language models LVLMs has been accompanied by advances in safety alignment mechanisms designed to prevent harmful content generation. However, these defenses remain vulnerable to sophisticated adversarial attacks. Existing jailbreak methods typically...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/07/29 12:0 a.m.4 views

Secure Tug-Of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security

The rapid advancement of multimodal large language models MLLMs has led to breakthroughs in various applications, yet their security remains a critical challenge. One pressing issue involves unsafe image-query pairs--jailbreak inputs specifically designed to bypass security constraints and elicit...

7.2AI score
SaveExploits0
Gitee
Gitee
added 2025/07/27 4:32 a.m.138 views

Exploit for Out-of-bounds Read in Openssl

This repository contains exploits and proof-of-concept vulnerability demonstration files from the team at Hacker House. The exploits target various vulnerabilities in different products and services, including: 1. AirWatch MDM solution: The repository contains a file called...

7.5CVSS9.3AI score0.99999EPSS
SaveExploits91
Packet Storm News
Packet Storm News
added 2025/07/17 12:0 a.m.7 views

Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility

AI systems are rapidly advancing in capability, and frontier model developers broadly acknowledge the need for safeguards against serious misuse. However, this paper demonstrates that fine-tuning, whether via open weights or closed fine-tuning APIs, can produce helpful-only models. In contrast to...

6.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/07/14 12:0 a.m.8 views

ARMOR: Aligning Secure and Safe Large Language Models Via Meticulous Reasoning

Large Language Models LLMs have demonstrated remarkable generative capabilities. However, their susceptibility to misuse has raised significant safety concerns. While post-training safety alignment methods have been widely adopted, LLMs remain vulnerable to malicious instructions that can bypass...

7.4AI score
SaveExploits0
Wallarm Lab
Wallarm Lab
added 2025/07/08 11:0 a.m.10 views

Inside the AI Threat Landscape: From Jailbreaks to Prompt Injections and Agentic AI Risks

AI has officially moved out of the novelty phase. What began with people messing around with LLM-powered GenAI tools for content creation has rapidly evolved into a complex web of agentic AI systems that form a critical part of the modern corporate landscape. However, this transformation has give...

8.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/07/08 12:0 a.m.60 views

CAVGAN: Unifying Jailbreak and Defense of LLMs Via Generative Adversarial Attacks on Their Internal Representations

Security alignment enables the Large Language Model LLM to gain the protection against malicious queries, but various jailbreak attack methods reveal the vulnerability of this security mechanism. Previous studies have isolated LLM jailbreak attacks and defenses. We analyze the security protection...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/07/06 12:0 a.m.6 views

Attention Slipping: a Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

As large language models LLMs become more integral to society and technology, ensuring their safety becomes essential. Jailbreak attacks exploit vulnerabilities to bypass safety guardrails, posing a significant threat. However, the mechanisms enabling these attacks are not well understood. In thi...

7.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/23 12:0 a.m.9 views

Security Assessment of DeepSeek and GPT Series Models against Jailbreak Attacks

The widespread deployment of large language models LLMs has raised critical concerns over their vulnerability to jailbreak attacks, i.e., adversarial prompts that bypass alignment mechanisms and elicit harmful or policy-violating outputs. While proprietary models like GPT-4 have undergone extensi...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/22 12:0 a.m.9 views

InfoFlood: Jailbreaking Large Language Models with Information Overload

Large Language Models LLMs have demonstrated remarkable capabilities across various domains. However, their potential to generate harmful responses has raised significant societal and regulatory concerns, especially when manipulated by adversarial techniques known as "jailbreak" attacks. Existing...

7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/22 12:0 a.m.7 views

SecurityLingua: Efficient Defense of LLM Jailbreak Attacks Via Security-Aware Prompt Compression

Large language models LLMs have achieved widespread adoption across numerous applications. However, many LLMs are vulnerable to malicious attacks even after safety alignment. These attacks typically bypass LLMs' safety guardrails by wrapping the original malicious instructions inside adversarial...

7.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/22 12:0 a.m.9 views

Investigating Vulnerabilities and Defenses against Audio-Visual Attacks: a Comprehensive Survey Emphasizing Multimodal Models

Multimodal large language models MLLMs, which bridge the gap between audio-visual and natural language processing, achieve state-of-the-art performance on several audio-visual tasks. Despite the superior performance of MLLMs, the scarcity of high-quality audio-visual training data and computation...

6.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/22 12:0 a.m.7 views

Universal Jailbreak Suffixes Are Strong Attention Hijackers

We study suffix-based jailbreaks$\unicodex2013$a powerful family of attacks against large language models LLMs that optimize adversarial suffixes to circumvent safety alignment. Focusing on the widely used foundational GCG attack Zou et al., 2023, we observe that suffixes vary in efficacy: some...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/22 12:0 a.m.10 views

Alphabet Index Mapping: Jailbreaking LLMs through Semantic Dissimilarity

Large Language Models LLMs have demonstrated remarkable capabilities, yet their susceptibility to adversarial attacks, particularly jailbreaking, poses significant safety and ethical concerns. While numerous jailbreak methods exist, many suffer from computational expense, high token usage, or...

7.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/21 12:0 a.m.7 views

From LLMs to MLLMs to Agents: a Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem

Large language models LLMs are rapidly evolving from single-modal systems to multimodal LLMs and intelligent agents, significantly expanding their capabilities while introducing increasingly severe security risks. This paper presents a systematic survey of the growing complexity of jailbreak...

6.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/19 12:0 a.m.8 views

Probing the Robustness of Large Language Models Safety to Latent Perturbations

Safety alignment is a key requirement for building reliable Artificial General Intelligence. Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models. We argue that this stems from the shallow nature of existing...

6.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/17 12:0 a.m.7 views

LLM Jailbreak Oracle

As large language models LLMs become increasingly deployed in safety-critical applications, the lack of systematic methods to assess their vulnerability to jailbreak attacks presents a critical security gap. We introduce the jailbreak oracle problem: given a model, prompt, and decoding strategy,...

7.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/12 12:0 a.m.6 views

SoK: Evaluating Jailbreak Guardrails for Large Language Models

Large Language Models LLMs have achieved remarkable progress, but their deployment has exposed critical vulnerabilities, particularly to jailbreak attacks that circumvent safety mechanisms. Guardrails--external defense mechanisms that monitor and control LLM interaction--have emerged as a promisi...

7.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/10 12:0 a.m.5 views

Evaluation Empirique De La Sécurisation Et De L'Alignement De ChatGPT Et Gemini: Analyse Comparative Des Vulnérabilités Par Expérimentations De Jailbreaks

Large Language models LLMs are transforming digital usage, particularly in text generation, image creation, information retrieval and code development. ChatGPT, launched by OpenAI in November 2022, quickly became a reference, prompting the emergence of competitors such as Google's Gemini. However...

7.3AI score
SaveExploits0
Rows per page
Query Builder