Lucene search
+L

545 matches found

Packet Storm News
Packet Storm News
•added 2026/09/01 12:00 a.m.•17 views

HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation

Production LLMs must handle inputs that attempt to override system instructions, bypass safety policies or elicit harmful responses. A common mitigation is a separate guardrail model. Existing reports, however, provide little evidence on Russian prompt injection or Russian surface obfuscation. We...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/08/31 12:00 a.m.•9 views

The Fragility of Jailbreak Robustness across Operational States

Existing jailbreak evaluations typically characterize robustness using a single attack success rate ASR measured in a default configuration the vanilla state. However, user-LLM interactions can induce diverse operational states beyond the vanilla state. In this work, we find that jailbreak...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/08/31 12:00 a.m.•10 views

EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities

Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-teaming methods treat this as a generation...

5.9AI score
SaveExploits0
GithubExploit
GithubExploit
•added 2026/08/30 2:30 p.m.•24 views

mo

PS5 UMTX Jailbreak The exploit code is largely based on the l...

5.9AI score
SaveExploits0
GithubExploit
GithubExploit
•added 2026/08/30 2:14 a.m.•54 views

ps4-web-exploit-host

PS4 9.00 jailbreak host GoldHEN, offline-caching Built by...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/08/27 12:00 a.m.•42 views

Circuit Discovery Helps Detect LLM Jailbreaking: A Mechanistic Interpretability Study

Despite extensive safety alignment, large language models LLMs remain vulnerable to jailbreak attacks that bypass safeguards to elicit harmful content. While prior work attributes this vulnerability to safety training limitations, the internal mechanisms by which LLMs process adversarial prompts...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/08/26 12:00 a.m.•18 views

SkillShield: Prompt-Space Security Skills for LLM Coding Agents

A coding agent edits files and executes shell commands with its developer's privileges, allowing malicious requests to translate directly into harmful actions or functional malware. Existing defenses have complementary limitations: weight-level alignment is unavailable to API-only deployers,...

6AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/08/26 12:00 a.m.•17 views

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

Large language models LLMs remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs. As jailbreak strategies keep emerging, defenses have proliferated in an ongoing cat-and-mouse gam...

5.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/08/26 12:00 a.m.•49 views

NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation

Safety evaluation is critical for assessing whether aligned Large Language Models LLMs remain robust against jailbreak attacks. Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt typically requires generating a target-model response to...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/08/26 12:00 a.m.•15 views

MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities

Multimodal Large Language Models MLLMs are increasingly deployed in real-world applications, yet how different factors shape their jailbreak vulnerabilities remains poorly understood. Existing benchmarks often couple harmful intent, prompt framing, visual semantics, and instruction carrier within...

5.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/08/18 12:00 a.m.•53 views

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate ASR without accounting for its dependence on attack budgets, resulting in unfair comparisons across methods. Existing compute-aware evaluations reduce heterogeneous...

5.5AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/08/17 8:23 a.m.•12 views

AI-Infra-Guard v4.5.2

📖 Documentation | 🌐 🇨🇳 中文 · 🇯🇵 日本語 · 🇪🇸 Español · 🇩🇪 Deutsch · 🇫🇷 Français · 🇰🇷 한국어 · 🇧🇷 Português · 🇷🇺 Русский 🚀 AI Red Teaming Platform by Tencent Zhuque Lab A.I.G AI-Infra-Guard integrates capabilities such as ClawScanOpenClaw Security Scan, Agent Scan,AI infra vulnerability scan, MCP Server &...

6.5AI score
SaveExploits0References33
Kitploit
Kitploit
•added 2026/08/17 4:19 a.m.•23 views

augustus v0.14.20

Augustus - LLM vulnerability scanner for prompt injection, jailbreak, and adversarial attack testing Augustus - LLM Vulnerability Scanner Test large language models against 210+ adversarial attacks covering prompt injection, jailbreaks, encoding exploits, and data extraction. Augustus is a Go-bas...

6.2AI score
SaveExploits0References9
Packet Storm News
Packet Storm News
•added 2026/08/10 12:00 a.m.•25 views

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the attacks that succeed. Harmful intent is a property of the prompt...

5.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/08/07 12:00 a.m.•12 views

The Anatomy of a Prompt Injection: A Component Model for Structured Analysis

Four years after prompt injection was first identified in 2022, attacks are still predominantly documented as verbatim strings rather than structured exploits, despite advancing agent capabilities and threat actors embedding injections to subvert AI-assisted security analysis. This paper formaliz...

9.3CVSS8.8AI score0.08034EPSS
SaveExploits1
Kitploit
Kitploit
•added 2026/08/05 1:46 p.m.•19 views

IOSSecuritySuite v2.3.0

⭐️ Do you want to become a certified iOS Application Security Engineer? ⭐️ Check out our practical & fully online course at: https://courses.securing.pl/courses/iase ISS Description by @r3ggi 🌏 iOS Security Suite is an advanced and easy-to-use platform security & anti-tampering library written in...

6.2AI score
SaveExploits0References26
Kitploit
Kitploit
•added 2026/08/05 8:29 a.m.•16 views

augustus v0.14.15

Augustus - LLM vulnerability scanner for prompt injection, jailbreak, and adversarial attack testing Augustus - LLM Vulnerability Scanner Test large language models against 210+ adversarial attacks covering prompt injection, jailbreaks, encoding exploits, and data extraction. Augustus is a Go-bas...

6.2AI score
SaveExploits0References9
Kitploit
Kitploit
•added 2026/08/05 6:21 a.m.•21 views

garak v0.16.0

garak, LLM vulnerability scanner Generative AI Red-teaming & Assessment Kit garak checks if an LLM can be made to fail in a way we don't want. garak probes for hallucination, data leakage, prompt injection, misinformation, toxicity generation, jailbreaks, and many other weaknesses. If you know nm...

6AI score
SaveExploits0References11
Packet Storm News
Packet Storm News
•added 2026/08/03 12:00 a.m.•69 views

AI Security Leaderboard: Methodology, Results and Minimal Standard

Frontier AI model developers increasingly rely on layered safeguards to prevent catastrophic misuse, but little public evidence exists on how much protection these safeguards provide, or how consistently across developers. We introduce the FAR.AI Minimal Standard for Safeguards, Version 1.0: a...

5.5AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/08/02 7:27 p.m.•14 views

augustus v0.14.13

Augustus - LLM vulnerability scanner for prompt injection, jailbreak, and adversarial attack testing Augustus - LLM Vulnerability Scanner Test large language models against 210+ adversarial attacks covering prompt injection, jailbreaks, encoding exploits, and data extraction. Augustus is a Go-bas...

6.2AI score
SaveExploits0References9
Rows per page
Query Builder