52 matches found
autoguardrails
autoguardrails Código abierto por Santander AI Lab. Una biblioteca / arnés de evaluación de investigación en seguridad de LLM / IA estilo autoresearch: busca sobre una única superficie mutable policy.md para minimizar la tasa de éxito de ataques ASR frente a un conjunto de evaluación fijo, con un...
Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models
Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with...
Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
Security benchmarks for LLM-based agents often report the attack success rate ASR as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences t...
Attack Success Rate Is Not a Number: On Measurement Validity in Agentic AI Security Evaluation
Attack success rate ASR is the headline metric in nearly every published evaluation of attacks on, and defenses for, LLM agents. We argue that ASR as currently used is not a single quantity but a family of metrics parameterized by six design choices that papers seldom specify and never hold...
Rethinking Backdoor Repair Evaluation: Distinguishing Aggregate Clean Utility from Benign Performance Preservation
Backdoor repair aims to suppress malicious behavior in compromised models while preserving benign task performance. Existing studies typically evaluate these objectives using Attack Success Rate ASR and Overall Clean Accuracy, but aggregate clean accuracy can obscure substantial degradation...
The Fragility of Jailbreak Robustness across Operational States
Existing jailbreak evaluations typically characterize robustness using a single attack success rate ASR measured in a default configuration the vanilla state. However, user-LLM interactions can induce diverse operational states beyond the vanilla state. In this work, we find that jailbreak...
Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning
Coding agents are increasingly used for software engineering tasks, including bootstrapping projects from third-party repositories whose integrity cannot be assumed. Prior work on repository poisoning largely focuses on attacker-controlled injection and disguise, but developers also shape risk...
Workspace Topology As an Attack Vector in Agentic Coding Assistants
Agentic coding assistants are finding widespread use, not just in new code development but in quickly ingesting and leveraging third-party code. This opens up a risk of malicious code being ingested as these coding tools operate with broad filesystem access inside developer workspaces. In this...
A Multimodal Automatic Redteaming Evaluation Based on Atomic Jailbreak Strategy Decoupling and Combination
Multimodal Large Language Models MLLMs have achieved impressive progress in image-text comprehension and generation, yet they remain susceptible to jailbreak attacks that can trigger harmful outputs and pose serious safety concerns. Existing multimodal jailbreak attacks have shown the feasibility...
Decoy Images Amplify Caption-Mediated Defenses against Encoded Jailbreaks
We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models VLMs: pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate ASR. The operative change is in the defense pipeline, not in the...
ALIBI: Adaptive Agentic Attacks on LLM-Based Vulnerability Detectors Via Adversarial Code Comments
Large language models are increasingly deployed for security-sensitive tasks such as vulnerability detection and code review. Their reliance on natural-language context embedded in source code exposes a previously underexplored attack surface: adversarial comments that can influence a detector's...
Defense against LLM Backdoors Using Critical Neuron Isolation Pruning
Large language models LLMs are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generally fall into inference-time detection or training-time mitigation, but face two key limitations. First, they focus on fine-tuning-based backdoors e.g., PEFT...
Understanding and Evaluating Claw-Like Agent Security through a Computer-Systems Lens
Claw-like AI agents e.g., OpenClaw are always-on processes with persistent access to credentials, files, tools, and external services. They take on system-level responsibilities -- installing packages, maintaining state, scheduling subtasks, and mediating I/O -- making security failures far more...
Model Poisoning against Federated Model Adaptation with Chain of Bit-Flips
Federated Learning FL allows a set of clients to collectively train a global model without sharing local training data. Giving the responsibility of the training to decentralized actors may lead to poisoning attacks: clients controlled by malicious third party potentially poison the training...
POISE: Position-Aware Undetectable Skill Injection on LLM Agents
Agent skills provide a lightweight mechanism for extending general-purpose agents, but their open format exposes them to skill-poisoning attacks. A practically dangerous injection must stay invisible: if executing the payload derails the user's legitimate task, the resulting failure signal invite...
MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models
Diffusion large language models dLLMs generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs. Because mask tokens are native inputs and tokens are committed by confidence rather than position,...
Defenses and Enablers for Skill Injection Attacks on Terminal Based Agents
Large language model LLM agents increasingly rely on reusable skills i.e. documents describing task-specific procedures. However, this introduces a new attack surface for agents to manage. We study two complementary directions for this threat. First, we evaluate guardian-based defenses: an...
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
As Large Language Models evolve for user convenience, vulnerability to jailbreak attacks continues to be reported despite ongoing efforts in safety training. Traditional jailbreak techniques typically focus on a single prompt injection, neglecting the models' ability to remember the flow of...
Adversarial Vulnerability under Temporal Concept Drift: A Longitudinal Study of Android Malware Detection
We present a longitudinal, drift-aware evaluation of adversarial robustness across more than a decade of Android applications using static and dynamic feature representations extracted from emulator and real-device executions. The dataset is organized into yearly slices and evaluated under three...
Evaluating Jailbreaking Vulnerabilities in LLMs Deployed As Assistants for Smart Grid Operations: A Benchmark against NERC Standards
The deployment of Large Language Models LLMs as assistants in electric grid operations promises to streamline compliance and decision-making but exposes new vulnerabilities to prompt-based adversarial attacks. This paper evaluates the risk of jailbreaking LLMs, i.e., circumventing safety alignmen...