1650 matches found
CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering
Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in...
Trajectory-Level Security Debt in LLM Coding Agents
LLM coding agents can traverse hundreds of intermediate code states before submitting a solution. Evaluating only the final artifact leaves the evolution of security findings unmeasured. We introduce the Security Debt Line Integral SDLI, which accumulates static-analysis risk when an agent reache...
"Nothing to See Here'': Unintended Disclosure through Revision Traces of LLM Deliverables
Large language model LLM assistants increasingly help users draft content for third-party recipients. During private drafting, the user or the model may introduce an item and later remove or replace it. The model may remove the item from the intended content but reveal it again when stating the...
Continuous Assurance of Agentic Security Auditors for Software Delivery Decision Gates
Large language model LLM-based repository auditors are increasingly deployed as security controls within continuous integration CI pipelines, where their findings admit, block, or delay software changes. As Agentic Software Development Life Cycle SDLC Security Controls, their non-deterministic...
CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-Based Agents
Large language model LLM-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection IPI. This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based...
Self-Evolving Defense: Continual Security Policy Learning for LLM Agents
Large language models LLMs increasingly power agents that access sensitive information, use external tools, and modify software repositories. Although these capabilities offer substantial benefits, they also create security risks such as jailbreaks, prompt injection, and vulnerable code generatio...
Share-Borne AI Virus: Memory-Hopping Attacks across LLM Agents
Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent...
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactl...
Xalgorix Autonomous AI Pentesting Agent 4.6.113
Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported - so you get proof, not a pile of maybes to triage. Self-hosted, private, and bring-your-own-LLM. Built in Go + TypeScript...
CyberPersistBench: Evaluating LLM-Based Cyber Attackers on Installation and Persistence
While LLM-based attackers exhibit growing proficiency in vulnerability exploitation, most existing cybersecurity benchmarks suffer from single-stage truncation, prematurely terminating evaluation upon initial access. In practice, initial footholds are exceptionally fragile across operational...
Separation of Duties for Privileged LLM Agents: A Governed Execution Architecture with Measured Security-Utility Trade-Offs
Large language model agents are increasingly granted real privileges executing commands, modifying files, calling APIs, so an agent that errs has already acted. Existing defences concentrate on the agent's inputs, while the path from a candidate action to privileged side effects remains less...
Inference-Layer Security: Defending against Adversarial Inference and Infrastructure Abuse
A Technical Report: Operating a large language model LLM as a service requires more than inference infrastructure: the provider must also defend against adversarial interactions that seek to exploit the service, including jailbreaking for harmful use, sophisticated denial of service, and...
ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability from Scratch
Large language model LLM agents are increasingly evaluated on cybersecurity tasks such as vulnerability reproduction, exploitation, and patching. However, existing cybersecurity benchmarks predominantly operate under a post-environment evaluation paradigm, i.e., handing the agent source code, a...
JEV As a Judge for Agent Trace Security: An Empirical Comparison with Generative LLM Judges
Security evaluation of tool-using agents requires judging actions in context, yet generative judges add latency, explanation overhead, and output-validation failures. We study whether JEV, a typed decision model, offers a useful alternative for retrospective trace classification. We evaluate JEV...
Know the Normal, Track the Attack: Context-Grounded and Stateful LLM Investigation over System Provenance
Provenance-based intrusion detection systems PIDSs identify suspicious activity in audit streams, but their outputs remain difficult to turn into coherent attack narratives. Direct LLM analyses of local anomalous subgraphs lack deployment-specific normal-behavior knowledge and validated attack...
Can Prompt Anonymity Protect Your Identity from LLM Providers?
User conversations with large language models LLMs often contain highly sensitive personal information that can be exploited by LLM providers to create detailed user dossiers, enable targeted advertising, and train more powerful models. To protect user privacy, anonymizing LLM proxies have emerge...
When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents
LLM agents increasingly rely on user approval to authorize security-sensitive actions at runtime. Such approvals are granted within a specific task and execution context. In long-lived agents, authorization decisions may need to persist across tasks or sessions. We find that this continuity can...
Climbing the Hill: Prompt Injection Red-Teaming against Frontier Models with Curriculum Reinforcement Learning
Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning RL to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such ...
HESP: Separating What to Probe from When to Stop in Local LLM Alert-Triage Agents
Security operations centers receive far more alerts than analysts can investigate, and organizations that cannot send their telemetry to hosted models must automate triage with small open-weight LLMs on their own hardware. Current LLM agents leave the investigation procedure to the model, and sma...
Xalgorix Autonomous AI Pentesting Agent 4.6.107
Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported - so you get proof, not a pile of maybes to triage. Self-hosted, private, and bring-your-own-LLM. Built in Go + TypeScript...