Lucene search
+L

107 matches found

Kitploit
Kitploit
added 2026/08/28 3:28 a.m.2 views

dataset

🚀 CySecBench: Dataset di prompt incentrati sulla cybersecurity basato su IA generativa per il benchmarking dei modelli linguistici di grandi dimensioni 🛡️ Il più grande e completo dataset di prompt incentrati sulla cybersecurity basato su IA generativa per il benchmarking dei modelli linguistici d...

5.8AI score
SaveExploits0References12
Kitploit
Kitploit
added 2026/08/27 7:30 p.m.1 views

PS4-5.05-Kernel-Exploit

PS4 5.05 Kernel-Exploit Zusammenfassung In diesem Projekt finden Sie eine vollständige Implementierung des zweiten „bpf"-Kernel-Exploits für die PlayStation 4 auf 5.05. Er ermöglicht es Ihnen, beliebigen Code als Kernel auszuführen, um Jailbreaking und Kernel-Level-Modifikationen am System zu...

5.4AI score
SaveExploits0References4
Kitploit
Kitploit
added 2026/08/26 9:09 p.m.4 views

ai-llm-red-team-handbook

Manuale Operativo per il Red Team AI/LLM e Manuale del Consulente Una toolkit operativa completa per condurre valutazioni di red team AI/LLM su modelli linguistici di grandi dimensioni, agenti AI, pipeline RAG e applicazioni abilitate all'AI. Questo repository fornisce sia indicazioni tattiche su...

5.3AI score
SaveExploits0References1
Kitploit
Kitploit
added 2026/08/21 5:28 p.m.3 views

PS4-5.05-Kernel-Exploit

Exploit del kernel PS4 5.05 Riepilogo In questo progetto troverai un'implementazione completa del secondo exploit del kernel "bpf" per PlayStation 4 sulla versione 5.05. Ti permetterà di eseguire codice arbitrario in modalità kernel, per consentire il jailbreak e modifiche a livello di kernel del...

5.4AI score
SaveExploits0References3
Packet Storm News
Packet Storm News
added 2026/06/05 12:00 a.m.28 views

Beyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (And Fail) Red Team Attacks

Standard AI red teaming evaluations reduce adversarial campaigns to a single binary outcome, attack success rate ASR, not taking into account the sequential structure of how models resist or yield to attacks. We propose applying process mining, a discipline for discovering and analyzing process...

5.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/06/04 12:00 a.m.115 views

Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense

Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifiers cannot adapt to these evolving attacks, while adaptive memory-based guardrails tend to over-refuse benign queries that resemble stored attacks...

5.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/06/01 12:00 a.m.14 views

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models

Diffusion large language models dLLMs generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs. Because mask tokens are native inputs and tokens are committed by confidence rather than position,...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/27 12:00 a.m.16 views

Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking

Jailbreak attacks on large language models LLMs aim to induce LLMs to produce content that they are expected to refuse. Automated black-box jailbreak generation is especially important for safety evaluation, where the attacker observes only model outputs and needs to automatically search for...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/26 12:00 a.m.32 views

Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security

Large Language Models LLMs are increasingly vulnerable to adversarial prompts that exploit semantic ambiguities to bypass safety mechanisms, resulting in harmful or inappropriate outputs. Such attacks, including jailbreaking and prompt injection, pose significant risks to the integrity and...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/11 12:00 a.m.20 views

Guaranteed Jailbreaking Defense Via Disrupt-And-Rectify Smoothing

This paper proposes a guaranteed defense method for large language models LLMs to safeguard against jailbreaking attacks. Drawing inspiration from the denoised-smoothing approach in the adversarial defense domain, we propose a novel smoothing-based defense method, termed Disrupt-and-Rectify...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/11 12:00 a.m.14 views

Re-Triggering Safeguards within LLMs for Jailbreak Detection

This paper proposes a jailbreaking prompt detection method for large language models LLMs to defend against jailbreak attacks. Although recent LLMs are equipped with built-in safeguards, it remains possible to craft jailbreaking prompts that bypass them. We argue that such jailbreaking prompts ar...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/11 12:00 a.m.16 views

LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content safety: behavior jailbreak, where an adversary induces an agent to execute dangerous OS-level operations with irreversible consequences. Existing...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/08 12:00 a.m.17 views

Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment

Recent advancements in visual context compression enable MLLMs to process ultra-long contexts efficiently by rendering text into images. However, we identify a critical vulnerability inherent to this paradigm: lowering image resolution inadvertently catalyzes jailbreaking. Our experiments reveal...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/04/27 12:00 a.m.11 views

Jailbreaking Frontier Foundation Models through Intention Deception

Large vision-language models exhibit remarkable capability but remain highly susceptible to jailbreaking. Existing safety training approaches aim to have the model learn a refusal boundary between safe and unsafe, based on the user's intent. It has been found that this binary training regime ofte...

5.3AI score
SaveExploits0
Schneier on Security
Schneier on Security
added 2026/03/10 9:50 a.m.15 views

Jailbreaking the F-35 Fighter Jet

Countries around the world are becoming increasingly concerned about their dependencies on the US. If you've purchase US-made F-35 fighter jets, you are dependent on the US for software maintenance. The Dutch Defense Secretary recently said that he could jailbreak the planes to accept third-party...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/03/06 12:00 a.m.8 views

Two Frames Matter: A Temporal Attack for Text-To-Video Model Jailbreaking

Recent text-to-video T2V models can synthesize complex videos from lightweight natural language prompts, raising urgent concerns about safety alignment in the event of misuse in the real world. Prior jailbreak attacks typically rewrite unsafe prompts into paraphrases that evade content filters...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/02/27 12:00 a.m.11 views

Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking

Jailbreak techniques for large language models LLMs evolve faster than benchmarks, making robustness estimates stale and difficult to compare across papers due to drift in datasets, harnesses, and judging protocols. We introduce JAILBREAK FOUNDRY JBF, a system that addresses this gap via a...

6AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/02/18 12:00 a.m.16 views

Recursive Language Models for Jailbreak Detection: A Procedural Defense for Tool-Augmented Agents

Jailbreak prompts are a practical and evolving threat to large language models LLMs, particularly in agentic systems that execute tools over untrusted content. Many attacks exploit long-context hiding, semantic camouflage, and lightweight obfuscations that can evade single-pass guardrails. We...

5.6AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/02/11 12:00 a.m.12 views

Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models

Jailbreaking large language models LLMs has emerged as a critical security challenge with the widespread deployment of conversational AI systems. Adversarial users exploit these models through carefully crafted prompts to elicit restricted or unsafe outputs, a phenomenon commonly referred to as...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/02/06 12:00 a.m.33 views

TrapSuffix: Proactive Defense against Adversarial Suffixes in Jailbreaking

Suffix-based jailbreak attacks append an adversarial suffix, i.e., a short token sequence, to steer aligned LLMs into unsafe outputs. Since suffixes are free-form text, they admit endlessly many surface forms, making jailbreak mitigation difficult. Most existing defenses depend on passive detecti...

5.3AI score
SaveExploits0
Rows per page
Query Builder