1655 matches found
LAAF: A Layered Accountability Architecture Framework for LLM Applications
Large Language Models LLMs operate in hospitals, courtrooms, banks, and public service desks, where fluent, confident outputs are treated as authoritative even when ungrounded or incorrect. When such an output contributes to harm, who is answerable, and through what mechanisms can responsibility ...
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent...
When Context Gets Root: Privilege Escalation in LLM Harnesses
Instruction hierarchy is a model-side defense that assigns instructions different levels of privilege according to their sources. These levels constrain which content may direct model behavior. During agent execution, however, agent harnesses construct context for each model invocation. This...
ContextLeak: Exfiltrating LLM Agent Context Via Malicious Tools
Exfiltrating an LLM agent's runtime context -- such as the user prompt, execution trajectory, and tool list -- poses severe security and privacy risks to users. Such attacks can be carried out via malicious tools and typically require three conditions: 1 the agent selects the malicious tool for...
Operationalizing Regulations into Code: A Model to Enhance Governance and Compliance in LLM Selection for Software Engineering
Integrating Large Language Models LLMs into the Software Development Life Cycle SDLC can improve developer productivity, but it also introduces security, privacy, and compliance risks during model selection. Regulations and frameworks such as the EU AI Act, the NIST AI Risk Management Framework...
OWASP Top 10 for LLM Applications 2026
OWASP Top 10 for LLM Applications 2026 is the latest community-driven guide to the most critical security risks facing applications powered by large language models. Developed by hundreds of AI security experts, this edition introduces updated rankings, expanded threat coverage, and new research...
Decoupling Is a Necessity: Transformation-Agnostic Decompiled Code Recovery under Optimization and Obfuscation
Reverse engineering is essential for software security analysis and vulnerability detection. Decompilation, the process of lifting binaries to high-level pseudocode, is central to this task. However, production binaries are hostile environments: aggressive compiler optimizations and adversarial...
SPA: Securing Persistent LLM Agents across Queries with Plan-First Information-Flow Control
Large language model LLM agents increasingly operate over untrusted webpages, documents, tools, and persistent states while exercising authority over security-sensitive resources. Existing defenses typically protect either planning or individual tool interactions, but persistent agents face a...
FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs
Evaluating the ability of large language models LLMs to discover software bugs is increasingly important. Existing benchmarks typically evaluate this capability by asking the model to generate a proof-of-concept input that triggers a predefined target vulnerability. However, this setup may overlo...
The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions
Agent guardrails are checks that approve or refuse each action before an LLM executes it. Sometimes they refuse requests that are genuinely safe. This over-safety blocks deployment when a guardrail refuses an authorized task. Evaluating over-safety is hard: at the boundary an authorized action...
Pentest-Swarm-AI
The first open-source pentesting tool built on a real swarm — no...
How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation
Capture-the-Flag CTF benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evaluations rely on shallow binary judgments or aggregate scores, overlooking the agent's trajectory to the flag. Consequently actual exploitation is conflated with...
Benchmarking Confidential Computing Performance on NVIDIA Blackwell GPUs
This paper measures the performance impact of running large language model inference and training inside a Trusted Execution Environment TEE on NVIDIA B200 GPUs, using Intel Trust Domain Extensions TDX confidential VMs together with NVIDIA Confidential Computing CC on Blackwell GPUs. The...
Unsaid, Unsafe? Implicit Security Obligations in LLM-Based RTL Code Generation
Large Language Models LLMs generate register-transfer-level RTL code with rapidly improving functional correctness. Security of LLM-generated code, however, has been studied mainly for software, where flaws can still be patched after deployment. Insecure RTL offers no such remedy once taped out...
SkillShield: Prompt-Space Security Skills for LLM Coding Agents
A coding agent edits files and executes shell commands with its developer's privileges, allowing malicious requests to translate directly into harmful actions or functional malware. Existing defenses have complementary limitations: weight-level alignment is unavailable to API-only deployers,...
A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks
Large language models LLMs remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs. As jailbreak strategies keep emerging, defenses have proliferated in an ongoing cat-and-mouse gam...
NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation
Safety evaluation is critical for assessing whether aligned Large Language Models LLMs remain robust against jailbreak attacks. Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt typically requires generating a target-model response to...
Answer Is Cheap, Show Me the Evidence! Augmenting Automated Vulnerability Assessment with Evidence
Software vulnerability SV assessment helps prioritize remediation by characterizing reported vulnerabilities. Existing automated methods predict assessment results from SV reports SVRs, but often overlook information in rich text, such as screenshots and code snippets, as well as contextual...
Improper Neutralization of Input Used for LLM Prompting
Overview strands-agents-tools is an A collection of specialized tools for Strands Agents Affected versions of this package are vulnerable to Improper Neutralization of Input Used for LLM Prompting via the pythonrepl function in src/strandstools/pythonrepl.py. An attacker can execute arbitrary...
CVE-2026-78379
Amazon Strands Agents Tools (prior to 0.8.5 ) is affected by an improper input neutralization flaw in its python_repl tool. A remote attacker can craft a prompt that forwards the non_interactive_mode keyword argument through the batch tool, thereby bypassing the human consent gate and achieving a...