545 matches found
HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation
Production LLMs must handle inputs that attempt to override system instructions, bypass safety policies or elicit harmful responses. A common mitigation is a separate guardrail model. Existing reports, however, provide little evidence on Russian prompt injection or Russian surface obfuscation. We...
The Fragility of Jailbreak Robustness across Operational States
Existing jailbreak evaluations typically characterize robustness using a single attack success rate ASR measured in a default configuration the vanilla state. However, user-LLM interactions can induce diverse operational states beyond the vanilla state. In this work, we find that jailbreak...
EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities
Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-teaming methods treat this as a generation...
mo
PS5 UMTX Jailbreak The exploit code is largely based on the l...
ps4-web-exploit-host
PS4 9.00 jailbreak host GoldHEN, offline-caching Built by...
Circuit Discovery Helps Detect LLM Jailbreaking: A Mechanistic Interpretability Study
Despite extensive safety alignment, large language models LLMs remain vulnerable to jailbreak attacks that bypass safeguards to elicit harmful content. While prior work attributes this vulnerability to safety training limitations, the internal mechanisms by which LLMs process adversarial prompts...
SkillShield: Prompt-Space Security Skills for LLM Coding Agents
A coding agent edits files and executes shell commands with its developer's privileges, allowing malicious requests to translate directly into harmful actions or functional malware. Existing defenses have complementary limitations: weight-level alignment is unavailable to API-only deployers,...
MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities
Multimodal Large Language Models MLLMs are increasingly deployed in real-world applications, yet how different factors shape their jailbreak vulnerabilities remains poorly understood. Existing benchmarks often couple harmful intent, prompt framing, visual semantics, and instruction carrier within...
A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks
Large language models LLMs remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs. As jailbreak strategies keep emerging, defenses have proliferated in an ongoing cat-and-mouse gam...
NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation
Safety evaluation is critical for assessing whether aligned Large Language Models LLMs remain robust against jailbreak attacks. Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt typically requires generating a target-model response to...
Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets
Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate ASR without accounting for its dependence on attack budgets, resulting in unfair comparisons across methods. Existing compute-aware evaluations reduce heterogeneous...
AI-Infra-Guard v4.5.2
📖 Documentation | 🌐 🇨🇳 中文 · 🇯🇵 日本語 · 🇪🇸 Español · 🇩🇪 Deutsch · 🇫🇷 Français · 🇰🇷 한국어 · 🇧🇷 Português · 🇷🇺 Русский 🚀 AI Red Teaming Platform by Tencent Zhuque Lab A.I.G AI-Infra-Guard integrates capabilities such as ClawScanOpenClaw Security Scan, Agent Scan,AI infra vulnerability scan, MCP Server &...
augustus v0.14.20
Augustus - LLM vulnerability scanner for prompt injection, jailbreak, and adversarial attack testing Augustus - LLM Vulnerability Scanner Test large language models against 210+ adversarial attacks covering prompt injection, jailbreaks, encoding exploits, and data extraction. Augustus is a Go-bas...
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks
Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the attacks that succeed. Harmful intent is a property of the prompt...
The Anatomy of a Prompt Injection: A Component Model for Structured Analysis
Four years after prompt injection was first identified in 2022, attacks are still predominantly documented as verbatim strings rather than structured exploits, despite advancing agent capabilities and threat actors embedding injections to subvert AI-assisted security analysis. This paper formaliz...
IOSSecuritySuite v2.3.0
⭐️ Do you want to become a certified iOS Application Security Engineer? ⭐️ Check out our practical & fully online course at: https://courses.securing.pl/courses/iase ISS Description by @r3ggi 🌏 iOS Security Suite is an advanced and easy-to-use platform security & anti-tampering library written in...
augustus v0.14.15
Augustus - LLM vulnerability scanner for prompt injection, jailbreak, and adversarial attack testing Augustus - LLM Vulnerability Scanner Test large language models against 210+ adversarial attacks covering prompt injection, jailbreaks, encoding exploits, and data extraction. Augustus is a Go-bas...
garak v0.16.0
garak, LLM vulnerability scanner Generative AI Red-teaming & Assessment Kit garak checks if an LLM can be made to fail in a way we don't want. garak probes for hallucination, data leakage, prompt injection, misinformation, toxicity generation, jailbreaks, and many other weaknesses. If you know nm...
AI Security Leaderboard: Methodology, Results and Minimal Standard
Frontier AI model developers increasingly rely on layered safeguards to prevent catastrophic misuse, but little public evidence exists on how much protection these safeguards provide, or how consistently across developers. We introduce the FAR.AI Minimal Standard for Safeguards, Version 1.0: a...
augustus v0.14.13
Augustus - LLM vulnerability scanner for prompt injection, jailbreak, and adversarial attack testing Augustus - LLM Vulnerability Scanner Test large language models against 210+ adversarial attacks covering prompt injection, jailbreaks, encoding exploits, and data extraction. Augustus is a Go-bas...