90 matches found
ClawGuard
ClawGuard 🛡️ ClawGuard 是一款面向自主 Agent 的安全防护工具包,专为 OpenClaw 及类似 LLM 驱动的 Agent 设计。随着 Agent 获得更多执行代码、访问 API 和管理文件的权限,ClawGuard 提供了必要的安全护栏,帮助降低安全风险。 📖 为什么选择 ClawGuard? 自主 Agent 功能多样,但同时也引入了独特的安全攻击向量: Prompt 注入 :恶意输入劫持 Agent 逻辑 权限提升 :Agent 执行未经授权的系统级操作 数据泄露 :意外泄露 PII(个人身份信息)或 API Key 给 LLM 提供商 资源耗尽...
opentaint
AI 时代的开源污点分析引擎 面向应用安全的形式化污点分析——发现 AST 模式匹配引擎遗漏的问题,让 LLM 代理将漏洞转化为规则,在两者都无法单独胜任的场景中实现规模化。 English | 简体中文 | 繁體中文 | 한국어 | Deutsch | Español | Français | Italiano | Dansk | 日本語 | Polski | Русский | Bosanski | العربية | Norsk | Svenska | Português Brasil | ไทย | Türkçe | Українська | বাংলা | हिन्दी |...
ConcoLLMic
ConcoLLMic: 基于智能体(Agent)的混合符号执行(Concolic Execution) 论文 : IEEE S&P 2026 ConcoLLMic 是首个由 LLM 智能体驱动的、语言与理论无关的混合符号执行器。与需要针对特定语言实现、且难以处理约束求解的传统符号执行工具不同,ConcoLLMic: 适用于 任何 编程语言和环境交互 — 支持 C、C++、Python、Java、...,甚至多语言系统;无需额外的环境建模,并能高效处理它们。 处理多样的约束理论 — 包括浮点运算、字符串、结构化数据和位级操作。 实现更优的覆盖率 — 在大约 4 小时内,ConcoLLMic...
AgentWatcher
AgentWatcher AgentWatcher es una defensa basada en detección contra la inyección indirecta de prompts en agentes LLM. Primero ejecuta atribución causal de contexto sobre el contexto no confiable para encontrar los contextos más influyentes, y luego aplica un LLM monitor que clasifica esos context...
skyvern
🐉 Automate Browser-based workflows using LLMs and Computer Vision 🐉 Skyvern automates browser-based workflows using LLMs and computer vision. It provides a Playwright-compatible SDK that adds AI functionality on top of playwright, as well as a no-code workflow builder to help both technical and...
anamnesis-release
Anamnesis: LLM 漏洞生成评估 该仓库包含用于研究 LLM 智能体如何根据漏洞报告在存在漏洞缓解措施的情况下生成漏洞利用的评估框架。给定一个错误报告和概念验证触发器,智能体会分析易受攻击的软件,并生成能够绕过各种安全缓解措施的有效漏洞。 在实验中,我以 QuickJS 中的一个零日漏洞为起点,然后让基于 Opus 4.5 和 GPT-5.2 构建的智能体生成漏洞。实验中我改变了启用的保护机制和漏洞利用的要求。Opus 4.5 解决了许多任务,而 GPT-5.2...
llm-agent-testbed
🛡️ Banco de Pruebas de Seguridad para Agentes LLM Banco de Pruebas Empírico de Vulnerabilidades y Defensas para Agentes LLM con Llamada a Herramientas Un banco de pruebas de seguridad disciplinado que evalúa si los agentes LLM equipados con herramientas pueden ser manipulados para realizar...
ShadowMem
ShadowMem: Protección de agentes LLM contra amenazas de horizonte largo mediante memoria sombra Este repositorio contiene la publicación oficial del código para el artículo Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory. ShadowMem es un marco defensivo que mantiene una...
ActGuard
ActGuard ActGuard es una defensa de auditoría de acciones previa a la ejecución contra la inyección indirecta de prompts en agentes LLM que utilizan herramientas. Este repositorio contiene la implementación final de ActGuard y el entorno de ejecución basado en AgentDojo necesario para evaluarla. ...
Defenses-for-Tool-Integrated-LLM
Defensas universales para agentes LLM integrados con herramientas contra ataques adversarios Este repositorio contiene el código y los experimentos de nuestro proyecto sobre la defensa de agentes de modelos de lenguaje de gran tamaño LLM integrados con herramientas contra ataques adversarios...
benign-instruction-bench
Aprobar el examen con el que se entrenó Reevaluando los detectores de inyección de prompts donde los agentes LLM realmente los utilizan: en las salidas de herramientas que un agente lee. Los equipos eligen detectores de inyección por sus puntuaciones en los benchmarks. Comprobamos si esas...
RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial...
Towards a Unified Misuse Monitoring Benchmark
LLM agents increasingly act in multi-actor environments, exposing them to misuse from multiple sources: decomposition attacks, where a harmful request is split into innocuous sub-requests, and prompt injection attacks, where a compromised tool delivers a malicious instruction. Existing evaluation...
Compromise Is Not Consequence: Evaluating Task-Scoped Authorization in LLM Agents with Paired Replay
A tool-using model can follow a malicious instruction even when its credentials are valid. We study whether task-scoped authorization contains the resulting tool execution. Our paired-replay testbed samples a model request once and submits the same action, resource, and arguments to broad bearer,...
SkillPoison: Progressive Skill Poisoning Via Successful Experiences
Self-improving LLM agents increasingly distill successful experiences into persistent, reusable skills. Existing skill attack methods corrupt this learning pipeline by injecting malicious triggers, behaviors, or false facts into individual experiences or extracted skills. However, such attacks ar...
Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training
Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on...
agentic-dm-gateway
Agentic DM Gateway Security control plane for LLM agents over private chat typically Discord DMs. It sits in front of your agent. It decides who may talk, whether the session is unlocked, whether the process is paused, and whether this message is safe enough to forward. Your model and tools stay...
VulValidate: Auditing Function-Level Vulnerability Labels with Executable Evidence
Reliable learning-based vulnerability detection requires high-quality labels, yet datasets built from vulnerability-fixing commits may label functions as vulnerable simply because they were changed by a security patch. We present VulValidate, a framework that uses LLM agents to coordinate dynamic...
APEX: Active Protection at Execution Boundaries for LLM Agents
Indirect prompt injection IPI hides adversarial instructions in content that large language model LLM agents read at runtime. As agents compose heterogeneous capability units, including Tools, MCP servers, and Skills, the carriers of injection multiply, and defenses built to recognize attack...
Prompted to Discriminate: Generalizing Malicious-Input Probes in the Wild
LLM agents increasingly rely on activation probes as runtime monitors for prompt injection, jailbreaks, and unsafe requests, reading the model's own hidden state to catch a harmful input before the agent acts on it. A cheap, increasingly common move, borrowed from LLM-as-judge prompting, is to...