1651 matches found
API Secrets Should Never Become Tokens in the LLM's Vocabulary
Tool-using large language model LLM agents turn credential hygiene from a storage problem into an execution-security problem. A key pasted into a prompt, or embedded in a system prompt or tool configuration, crosses from an authentication boundary into a data pipeline, where it may persist in...
Climbing the Hill: Prompt Injection Red-Teaming against Frontier Models with Curriculum Reinforcement Learning
Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning RL to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such ...
Xalgorix Autonomous AI Pentesting Agent 4.6.99
Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported - so you get proof, not a pile of maybes to triage. Self-hosted, private, and bring-your-own-LLM. Built in Go + TypeScript...
ORBIT: A Framework for Multi-Agent Safety and Security Evaluations
Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet they give rise to significant safety and security risks: the flexible protocols that enable task generalization also expose novel threats, fr...
ethibench
From Controlled to the Wild: Evaluation of Pentesting Agents in the Real-World AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which systems will perform best on real-world targets. Most existing evaluations...
MetaPermit: Scalable and Auditable Access Control for AI Agents Via LLM-Inferred Meta-Attributes
The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection IPI attacks. In practice, deployed agent systems such as OpenAI Codex and Claude Code protect tool...
Configuration, Not Conscience: A Large-Scale Empirical Study of LLM System Prompts
Leaked system prompts are often treated as windows into the hidden values of commercial language models, yet their composition is rarely studied at scale. We analyze a merged corpus of 407 leaked, reconstructed, or officially published system prompts from 62 vendors across four community...
AuthGuard-R: Safety-Compliant Mission Hijacking and Dual-Gate Defense for LLM-Controlled Robots
Large language models are increasingly used as high-level planners for mobile robots, robot manipulators, and autonomous vehicles. Recent studies show that these systems can be influenced through malicious text, speech, visual instructions, retrieved documents, and poisoned sensory context. Most...
Resource-Optimized and Energy-Aware Agentic AI Framework Anchored on Blockchain for Secure Software Supply Chains
This paper proposes a blockchain-backed agentic security framework designed to safeguard the complete software development lifecycle SDLC while also securing the agentic AI components responsible for monitoring it. The framework coordinates a set of specialised security agents, covering source...
AGATE: Provenance-Based Runtime Defense against Compositional Attacks on LLM Agents
LLM agents can produce harmful effects through sequences of ordinary operations. Judging such actions requires establishing both the authority that permits them and the origin of the data they carry. We present AGATE, an authorization and data-provenance gate at instrumented agent-harness...
Xalgorix Autonomous AI Pentesting Agent 4.6.95
Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported - so you get proof, not a pile of maybes to triage. Self-hosted, private, and bring-your-own-LLM. Built in Go + TypeScript...
FragToken: Amplifying LLM Inference Costs through Noncanonical Token Generation
As large language model LLM inference becomes increasingly expensive, resource-consumption attacks pose a growing threat to model providers. Existing attacks typically amplify cost by inducing abnormally long or repetitive outputs on attacker-controlled or triggered requests, making them easier t...
AI-powered fuzzing with the GitHub Security Lab Taskflow Agent
If you 're new to fuzzing and want to learn the fundamentals first, check out our Fuzzing 101 course at gh.io/fuzzing101. Continuous fuzzing is not a magic solution that solves all your problems . Even projects that have been enrolled in OSS-Fuzz for years can still hide critical bugs, and the...
CVE-2026-61732
Decepticon is an autonomous hacking agent for red teams. Versions prior to 1.1.17 wrap web crawl results — the output of agent reconnaissance against target services — into LLM messages without neutralizing ChatML special-token literals. Under the BYOK Bring Your Own Key deployment model, users...
EUVD-2026-86329
Decepticon is an autonomous hacking agent for red teams. Versions prior to 1.1.17 wrap web crawl results — the output of agent reconnaissance against target services — into LLM messages without neutralizing ChatML special-token literals. Under the BYOK Bring Your Own Key deployment model, users...
PT-2026-98224
Name of the Vulnerable Software and Affected Versions TREK versions prior to 3.3.0 Description When the LLM PARSING feature is enabled, an authenticated user with write permission to a trip instance can store an attacker-controlled llm base url via the settings API. This value is processed by the...
Instrumental Monitor Evasion Emerges under Ordinary Task Pressure
A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals. We study instrumental evasion, the propensity of LLM agents to circumvent runtime monitoring as a means of completing ordinary tasks. We introduce EvasionBench, a...
ClaimMirage: When Self-Claims in Domain Names Change LLM Threat Judgments
Short claims such as not-phishing or official can change how a large language model LLM judges a domain name, without explicit prompt-injection commands. We study this manipulation as ClaimMirage: a name under inspection claims its own safety or approval. We analyze 622,080 judgments across 64...
OllamaDrama: Designing and Deploying a Honeypot to Measure Attacks on Exposed LLM Infrastructure
Publicly exposed large language model LLM infrastructure creates a growing attack surface, yet real-world targeting remains poorly understood. We present Ollure, a low- and medium-interaction honeypot that emulates the Ollama API without a backend LLM. Spanning four deployments across cloud and...
OWASP LLM Top 10 2026: Every Move Points the Same Direction
This is Post 4 of a four-part series: Post 1: Agentic AI Security: The Chatbot Era Is Over Post 2: Generative AI Security: Why AI Needs a New Kind of Security Post 3: The OWASP LLM Top 10 Was the Warm-Up: What Comes Next The 2026 edition of the the OWASP Top 10 for LLM Applications, published by...