1650 matches found
trustmebro
Bypass llm guardrails by confusing it with fabricated tool output. Results · Installation · Quick start · Rules · Architecture TrustMeBro intercepts command-line tools invoked by coding agents such as Codex, Claude Code, and pi. Rules decide whether to return fabricated output, modify the real...
pentestcode
PentestCode 터미널에서 실행되는 AI 침투 테스트 에이전트. 멀티 에이전트 아키텍처 • 인게이지먼트 상태 추적 • 20개 이상의 LLM 제공자 PentestCode는 터미널에서 동작하는 자율 침투 테스트 에이전트입니다. 타겟을 지정하면 도구를 실행하고, 출력을 읽고, 네트워크에 대한 그림을 갱신하며, 다음에 무엇을 할지 스스로 결정합니다 — 마치 실제 운영자처럼요. OpenCodeMIT의 하드 포크로, 코드 편집 중심 기능을 걷어내고 공격적 보안을 위해 재구축했습니다. 베타 — 실제 인게이지먼트와 CTF에서 검증되었지...
COPEX: Benchmarking LLM Robustness to Adversarial Context across Model Context Protocol Layers
Large language models increasingly mediate tool use in Model Context Protocol MCP systems, where adversarial influence may enter through user instructions, tool schemas, tool outputs, or protocol messages. Existing benchmarks often evaluate deployed agents, conflating model susceptibility with...
The Same Zero: Why Identical ASR Can Imply Different Guarantees in LLM-Agent Security
LLM-agent security has produced a dense landscape of defenses - prompt hardening, content filters, permission gates, sandboxes - yet no framework tells a deployer what a defense actually guarantees, or where that guarantee comes from. We apply Verification Autonomy Levels VAL - L0: LLM...
Xalgorix Autonomous AI Pentesting Agent 4.6.139
Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported - so you get proof, not a pile of maybes to triage. Self-hosted, private, and bring-your-own-LLM. Built in Go + TypeScript...
APEX: Active Protection at Execution Boundaries for LLM Agents
Indirect prompt injection IPI hides adversarial instructions in content that large language model LLM agents read at runtime. As agents compose heterogeneous capability units, including Tools, MCP servers, and Skills, the carriers of injection multiply, and defenses built to recognize attack...
skyvern v1.0.55
🐉 Automate Browser-based workflows using LLMs and Computer Vision 🐉 Skyvern automates browser-based workflows using LLMs and computer vision. It provides a Playwright-compatible SDK that adds AI functionality on top of playwright, as well as a no-code workflow builder to help both technical and...
PT-2026-104401
crmne/ruby llm at commit fa6f279847d6d7027814539d9c0dfc3bbdfd2a83 contains a polynomial-time regular expression denial-of-service condition in Mistral model capability matching on Ruby 3.1.x...
Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents
Large language model LLM-based agents increasingly connect model-generated decisions to security-sensitive software capabilities such as command execution, filesystem access, network communication, browser control, and external tools. Existing analyses often use predefined sensitive operations as...
PrivDev: Mapping Static-Analysis Data Types to DPV
Static-analysis scanners can identify personal-data types in source code, but they lack mechanisms to connect these findings to standardized privacy vocabularies. PrivDev maps 122 Bearer CLI data types to Data Privacy Vocabulary Personal Data DPV-PD categories and links them to potentially releva...
Defense-In-Depth at the Perception-Reasoning Interface of LLM-Centric Agentic UAV Swarms
Large Language Models LLMs increasingly support Uncrewed Aerial Vehicle UAV swarm operations such as data collection scheduling, where the model reads structured sensor reports and decides which sensors to visit. An adversary who quietly manipulates those reports can redirect the swarm without...
Xalgorix Autonomous AI Pentesting Agent 4.6.136
Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported - so you get proof, not a pile of maybes to triage. Self-hosted, private, and bring-your-own-LLM. Built in Go + TypeScript...
Securing Computer-Use Agents against Branch Steering Attacks
Modern Computer Use Agents CUAs directly interact with graphical user interfaces and execute third-party web tools, exposing them to indirect prompt injection across every rendered page and tool response. While the Dual-LLM pattern is the primary system-level architecture offering formal security...
Decepticon: Role-boundary forgery via ChatML special-token literals in web crawl output composed into LLM context
SummaryDecepticon wraps web crawl results — the output of agent reconnaissance against target services — into LLM messages without neutralizing ChatML special-token literals. Under the BYOK Bring Your Own Key deployment model, users configure their own LLM credentials to any OpenAI-compatible...
MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication
Inter-agent communication is central to Large Language Model Multi-Agent Systems LLM-MAS, but it introduces an underexplored vulnerability: Agent-in-the-Middle AiTM attacks that manipulate messages in transit without compromising the agents themselves. Prior work reports Attack Success Rates ASR...
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable...
Prompted to Discriminate: Generalizing Malicious-Input Probes in the Wild
LLM agents increasingly rely on activation probes as runtime monitors for prompt injection, jailbreaks, and unsafe requests, reading the model's own hidden state to catch a harmful input before the agent acts on it. A cheap, increasingly common move, borrowed from LLM-as-judge prompting, is to...
Xalgorix Autonomous AI Pentesting Agent 4.6.128
Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported - so you get proof, not a pile of maybes to triage. Self-hosted, private, and bring-your-own-LLM. Built in Go + TypeScript...
Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations
Large language models LLMs in production systems face prompt injections, trojans backdoors, and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose Rstabf, ...
OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents
LLM agents with tool-calling capabilities can access external services and private user data, but they may retrieve more information than a user's request explicitly requires. We study this behavior in structured tool-calling agents and term it proactive over-authorization. This setting differs...