3791 matches found
CVE-2026-87912: Unverified Ownership
A missing S3 bucket ownership verification in the AWS Security Agent plugin in Amazon aws-agents-for-devsecops before 1.1.0 might allow remote attackers to obtain the private source archive of a scanned workspace, including credentials and infrastructure state contained in that archive, via a...
PT-2026-89469
Name of the Vulnerable Software and Affected Versions Amazon aws-agents-for-devsecops versions prior to 1.1.0 Description The AWS Security Agent plugin fails to verify the ownership of S3 buckets. This allows remote attackers to obtain the private source archive of a scanned workspace, which may...
CVE-2026-88062: Improper Control of Generation of Code
OmniRoute is an open-source AI gateway providing a single endpoint for multiple model providers. In 3.8.49 and earlier, the OmniRoute POST /api/acp/agents custom ACP agent endpoint accepted attacker-controlled binary and versionCommand values and used only a self-consistency check before...
HOL Guard AI Agent Runtime Security Benchmark
HOL Guard Benchmark is an open source fixture-based benchmark for evaluating security controls applied to AI coding agents. It defines deterministic scenarios covering secret access, risky shell execution, tool poisoning, approval handling, safe operations, and receipt generation across Codex CLI...
halo-record v0.2.42
halo-record Tamper-evident audit trails for AI agents — hash-chained Runtime Records, rendered as a Runtime Report your customers can check themselves. Every action your instrumentation captures tool calls, model calls, data access, approvals becomes one Runtime Record in an append-only,...
BIT-KIBANA-2026-78583 Incorrect Authorization in Kibana Leading to Privilege Escalation
Incorrect Authorization CWE-863 in Kibana can lead to privilege escalation via Input Data Manipulation CAPEC-153. Elasticsearch cluster privilege declarations originating from integration packages were not validated before being used to mint credentials for enrolled Elastic Agents. A user holding...
CyberStrikeAI v1.7.18
CyberStrikeAI 中文 | English The system of action for AI-native cybersecurity—where intent becomes governed execution, evidence becomes operational memory, and every operation improves the next. CyberStrikeAI connects planning, execution, human oversight, evidence, and replay in one auditable...
Flowise CSV Agent Prompt Injection Remote Code Execution Vulnerability
This vulnerability allows remote attackers to execute arbitrary code on affected installations of Flowise. Authentication is not required to exploit this vulnerability. The specific flaw exists within the run method of the CSVAgents class. The issue results from insufficient input sanitization wh...
AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents
Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion AgentBench or security robustness AgentDojo, ASB, rather than the complete pipeline of planning, tool selection, tool execution, memory and reasoning. Failures can occur at any stage, yet existing...
Big Enough to Break Out: Tracking the Rising Capability of LLM Penetration-Testing Agents
Large language model LLM agents are increasingly applied to penetration testing, but we still know little about what they can do or how they fail. We compare two PentestGPT-based systems: a legacy human-in-the-loop system running the open-weight Kimi K2.5, and a newer autonomous system running...
AIs as Modern Genies
This essay was written with Barath Raghavan, and originally appeared inLawfare. In April, an artificial intelligence AI agent conducting a routine task at a company hit a snag, tried to solve it, and soon ended up deleting the company’s database along with all of its backups. In July, OpenAI aske...
Autonomous AI Agents Compromise Thousands of Credentials in Under Six Hours
Threat actors are continuing to leverage artificial intelligence AI to streamline their operations, with one financially motivated hacking group employing an autonomous, multi-agent attack framework to carry out a large-scale credential harvesting campaign within six hours. Google Threat...
Dark-Moon v1.4.0
DarkMoon Open-source autonomous AI penetration testing — finds, exploits and qualifies every vulnerability on your own infrastructure ⭐ Star DarkMoon · 🚀 Quick Start · 🧩 Integrations · 📊 Benchmark · 🔒 Darkmoon Pro · ▶️ Install & Run OSS · ▶️ Demo Pro DarkMoon is autonomous AI penetration testing...
Flounder 0.4.0
Flounder turns modern coding agents into an end-to-end security audit system. Give it an authorized target boundary - a repository, source tree, package, deployed clue, or prior run - and the agent can prepare the workspace, read the code and supporting material, map the attack surface, dig into...
Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Memory Systems
Long-running language-model agents depend on persistent memory. Many agent-memory systems preserve history through soft revocation: a contradicted fact is marked invalid and retained rather than deleted. However, whether that mark is enforced at retrieval time is unexamined. In this paper, we...
VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities
The software supply chain has become an increasingly exposed attack surface because of its reliance on intricate yet fragile dependencies. Existing defenses such as GitHub Dependabot often raise many false alerts because their coarse-grained matching cannot determine whether a vulnerable dependen...
LLM-Based Penetration Testing in the Presence of Honeypots
Large language model LLM agents are increasingly employed for offensive cybersecurity tasks such as automated vulnerability discovery, reconnaissance, and penetration testing. This new capability also threatens one of the defender's most valuable tools: deception. Traditional honeypots rely on...
TrojanWorld: Backdooring World-Model Agents Via Imagination Steering
World models increasingly serve as the predictive core of model-based reinforcement learning agents, enabling them to simulate future dynamics and reason over imagined trajectories before acting. Their substantial training demands make pretrained world models attractive for distribution and reuse...
AgentDrift: A Step-Labeled Benchmark of Injection-Hijacked LLM Agent Trajectories
LLM agents complete tasks by issuing sequences of tool calls, and every observation they read is a channel through which an indirect prompt injection can enter. A successful injection has a characteristic shape when the trajectory is read in order: a benign prefix gives way to actions that serve...
OpenAI Agents Hacked Another Website
Plus: Tens of millions of US and Canadian drivers’ licenses go up for sale on the dark web, the US military finally tries to tackle the risk online ad data poses to troops, and more...