Lucene search
+L

273 matches found

Kitploit
Kitploit
•added 2026/07/28 10:11 p.m.•10 views

nuguard v0.8.8

NuGuard Open Source NuGuard is an open source AI application security toolkit. Its goal is to provide the most extensive redteaming and behavioral validation of Agentic AI applications. With NuGuard, AI developers can focus on building their applications while NuGuard continuously tests and...

5.9AI score
SaveExploits0References8
Kitploit
Kitploit
•added 2026/07/23 8:44 a.m.•9 views

afrog v3.5.6

A Security Tool for Bug Bounty, Pentest and Red Teaming English • 中文 Download • Wiki • Afrog PoC 规则编写权威指南 PoC Contributors 不动明王 | 雪山 | White-hua | 123456 | ifofor | Air | !Typora-Logoh...

6.6AI score
SaveExploits0References42
Kitploit
Kitploit
•added 2026/07/20 9:21 p.m.•14 views

promptfoo v0.121.19

Promptfoo: LLM evals & red teaming promptfoo is a CLI and library for evaluating and red-teaming LLM apps. Stop the trial-and-error approach - start shipping secure, reliable AI apps. Website · Getting Started · Red Teaming · Documentation · Discord Promptfoo is now part of OpenAI. Promptfoo...

5.8AI score
SaveExploits0References4
Packet Storm News
Packet Storm News
•added 2026/07/19 12:00 a.m.•36 views

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions

Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate ASR, yet an attack that breaks a model is not necessarily useful for improving its safety. We propose a defender-centric view of jailbreak evaluation, where attacks are evaluat...

5.4AI score
SaveExploits0
The Hacker News
The Hacker News
•added 2026/07/16 8:42 a.m.•33 views

OpenAI’s GPT-Red Automates Prompt Injection Testing to Harden GPT-5.6 Sol

OpenAI has disclosed details of GPT-Red , an internal automated red-teaming model that scales prompt injection vulnerability discovery with an aim to fix issues before the tools are deployed widely. "GPT‑Red is a strong red-teamer, and our previous models are highly vulnerable to its prompt...

6.1AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/07/15 5:13 p.m.•15 views

afrog v3.5.5

A Security Tool for Bug Bounty, Pentest and Red Teaming English • 中文 Download • Wiki • Afrog PoC 规则编写权威指南 PoC Contributors 不动明王 | 雪山 | White-hua | 123456 | ifofor | Air | !Typora-Logoh...

6.6AI score
SaveExploits0References42
Packet Storm News
Packet Storm News
•added 2026/07/13 12:00 a.m.•11 views

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation

Safety evaluation of large language models LLMs relies largely on single-turn attack datasets and single-judge scoring, underestimating risk from adaptive multi-turn adversaries and reporting a single success rate that does not separate partially actionable outputs from those carrying complete...

6.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/07/13 12:00 a.m.•11 views

Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming

Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety failures directly actionable. Red-teaming must therefore keep pace with evolving models and tools. Existing approaches mainly optimize attack success and preserv...

6.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/07/03 12:00 a.m.•12 views

CONTRA: Red-Teaming Configurations of Personalizable Agents

Recent tools such as OpenClaw have extended the capabilities of LLM-based agents from simple dialog-based systems to fully autonomous agents. These systems allow personalization of the agent through modifiable internal files and the installation of skills. While this enables deployment in a wide...

6.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/06/30 12:00 a.m.•13 views

Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming

The fast growth of open-source AI infrastructure, from model serving engines and agent platforms to the Model Context Protocol MCP ecosystem and the language models themselves, has outpaced the security tooling available to defend it. We present AI-Infra-Guard, an open-source framework that...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/06/25 12:00 a.m.•51 views

MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

Multimodal agentic retrieval-augmented generation RAG systems expand the attack surface beyond prompt injection to include text poisoning, image injection, direct-query attacks, and orchestrator-level tool manipulation. Existing red-teaming approaches are typically surface-specific and often...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/06/22 12:00 a.m.•14 views

RIFT-Bench: Dynamic Red-Teaming for Agentic AI Systems

Agentic AI systems powered by large language models LLMs are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities. Existing security evaluations are often tied to specific implementations or domains, limiting unified...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/06/18 12:00 a.m.•127 views

LLM Agent Safety, Multi-Turn Red-Teaming, Jailbreak Benchmarks, Adversarial Robustness, Safety-Critical Systems

Large language model LLM agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized. We present NRT-Bench, a benchmark for multi-turn red-teaming of LLM agents acting as...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/06/11 12:00 a.m.•39 views

MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems

Hierarchical multi-agent systems MAS are rapidly being deployed in high-stakes workflows across domains such as finance and software engineering. In these systems, safety and security are inherently distributed across role-specialized agents, significantly expanding the attack surface, particular...

5.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/06/10 12:00 a.m.•32 views

PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

Large Language Models LLMs are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted external sources. Existing defenses mainly focus on blocking malicious content at...

5.4AI score
SaveExploits0
Microsoft Secure
Microsoft Secure
•added 2026/06/04 7:14 p.m.•41 views

Updating the taxonomy of failure modes in agentic AI systems: What a year of red teaming taught us

In this article 1. Why the Taxonomy Needed Updating 2. Seven new failure modes 3. Operational findings: What red teaming showed 4. New mitigations 5. What to do this quarter When the Microsoft AI Red Team published the Taxonomy of Failure Modes in Agentic AI Systems in April 2025, the goal was a...

8.8CVSS7.4AI score0.24429EPSS
SaveExploits7
Packet Storm News
Packet Storm News
•added 2026/06/04 12:00 a.m.•47 views

RedEdit: Agentic Red-Teaming of Image Safety Classifiers Via MCTS-Guided Photo-Editing

Image safety classifiers serve as a critical component of contemporary content moderation systems on the internet. However, their resilience against user-style malicious image editing remains underexplored. Such behaviors are highly prevalent in daily scenarios but difficult to fully reproduce. T...

5.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/06/02 12:00 a.m.•24 views

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

Recent computer-using-agent CUA red-teaming papers report prompt-injection attack success rates ASR of 42-98%, but these headline numbers cluster on retired models and on the most-vulnerable model in each paper's panel. We ask whether those techniques, reproduced as hand-crafted templates, still...

5.5AI score
SaveExploits0
GithubExploit
GithubExploit
•added 2026/06/01 10:12 a.m.•120 views

-cascade-scan

cascade-scan AI Agent security evaluation framework — autom...

6.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/05/17 12:00 a.m.•65 views

ADR: An Agentic Detection System for Enterprise Agentic AI Security

We present the Agentic AI Detection and Response ADR system, the first large-scale, production-proven enterprise framework for securing AI agents operating through the Model Context Protocol MCP. We identify three persistent challenges in this domain: 1 limited observability -- existing Endpoint...

5.8AI score
SaveExploits0
Rows per page
Query Builder