6 matches found
IPI-exposure-signal
IPI Exposure Signal Este es el repositorio de código de nuestro artículo: Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure. Versión en arXiv y enlace al artículo: https://arxiv.org/abs/2608.02657 Este repositorio implementa el pipeline de sondeo para señales...
benign-instruction-bench
Aprobar el examen con el que se entrenó Reevaluando los detectores de inyección de prompts donde los agentes LLM realmente los utilizan: en las salidas de herramientas que un agente lee. Los equipos eligen detectores de inyección por sus puntuaciones en los benchmarks. Comprobamos si esas...
The Same Zero: Why Identical ASR Can Imply Different Guarantees in LLM-Agent Security
LLM-agent security has produced a dense landscape of defenses - prompt hardening, content filters, permission gates, sandboxes - yet no framework tells a deployer what a defense actually guarantees, or where that guarantee comes from. We apply Verification Autonomy Levels VAL - L0: LLM...
ActGuard
ActGuard ActGuard es una defensa de auditoría de acciones previa a la ejecución contra la inyección indirecta de prompts en agentes LLM que utilizan herramientas. Este repositorio contiene la implementación final de ActGuard y el entorno de ejecución basado en AgentDojo necesario para evaluarla. ...
Optimizing Agent Planning for Security and Autonomy
Indirect prompt injection attacks threaten AI agents that execute consequential actions, motivating deterministic system-level defenses. Such defenses can provably block unsafe actions by enforcing confidentiality and integrity policies, but currently appear costly: they reduce task completion...
Progent: Programmable Privilege Control for LLM Agents
LLM agents are an emerging form of AI systems where large language models LLMs serve as the central component, utilizing a diverse set of tools to complete user-assigned tasks. Despite their great potential, LLM agents pose significant security risks. When interacting with the external world, the...