31 matches found
xalgorix
Xalgorix — Open-source AI pentester that proves vulnerabilities Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported — so you get proof, not a pile of maybes to triage...
osmedeus
Osmedeus Osmedeus - A Modern Orchestration Engine for Security What is Osmedeus? Osmedeus is a security focused declarative orchestration engine that simplifies complex workflow automation into auditable YAML definitions, complete with encrypted data handling, secure credential management, and...
xalgorix
Xalgorix — Open-source AI pentester that proves vulnerabilities Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported — so you get proof, not a pile of maybes to triage...
BoxPwnr
BoxPwnr A fun experiment to see how far Large Language Models LLMs can go in solving CTF challenges and security labs on their own. It started with HackTheBox and now covers many platforms and agentic solvers. BoxPwnr provides a plug and play system that can be used to test performance of differe...
BoxPwnr-Traces
BoxPwnr-Traces BoxPwnr traces and benchmark results across multiple security platforms. Each trace includes the full LLM interaction, commands executed, a markdown report + attack graph, stats and config used. Browse leaderboards, replay runs in an interactive web viewer, and read AI-generated...
mcpguard-dynamic
MCPGuard-Dynamic Kernel-level sandboxing for LLM agent tool calls made through the Model Context Protocol MCP. MCPGuard sits as a transparent proxy between an MCP client the agent / runner and an MCP server subprocess, applying three layered defenses to every tool invocation. The lowest layer is...
auto-re-agent
re-agent Autonomous reverse-engineering agent — source-aware reverser/checker loop, objective verifier, parity engine, and Ghidra backend. Overview Demo: YouTube re-agent automates a reverse-engineering workflow by combining a reverser/checker loop with Ghidra decompilation through...
SWE-agent
!warning La mayor parte de nuestro esfuerzo actual de desarrollo está en mini-swe-agent, que ha reemplazado a SWE-agent. Iguala el rendimiento de SWE-agent, siendo mucho más simple. Consulta las FAQ para más detalles sobre las diferencias. Nuestra recomendación general es usar mini-SWE-agent en...
ROPE
ROPE: Routed Origin Policy Enforcement Source code of our paper: ROPE: Routed Origin Policy Enforcement against Indirect Prompt Injection by Xinhang Ma, Chaowei Xiao, William Yeoh, Ning Zhang, Yevgeniy Vorobeychik Abstract Indirect prompt injection IPI plants instructions in the content a...
cve-bench
CVE-Bench A benchmark for evaluating LLM agents on fixing real-world security vulnerabilities. Agents run inside sandboxed Docker containers and are scored against the maintainer's security test suite. Requirements Python 3.12+ Docker OPENAIAPIKEY, ANTHROPICAPIKEY, and/or POOLSIDEAPIKEY in your...
pike-agent
Pike Agent pike-agent records and analyzes how programs behave on Linux. It traces a program's activity, indexes it into a database, and lets you chat with an LLM agent about it in a TUI. Example of prompts: Crash diagnosis: This program crashed with a bus error. What happened? Race condition...
SecurityClaw
SecurityClaw — 自律SOCエージェントフレームワーク モジュール化されたスキルベースの自律型セキュリティ運用センター(SOC)エージェントです。OpenSearch/Elasticsearchのデータを監視し、RAGベースの行動メモリを構築し、LLMを使用してリアルタイムの異常を検証します。 Features スキルのモジュール性 — 機能を独立したフォルダとして管理。logic.py(Python)+ instruction.md(LLMガイダンス) ハートビートループ — Cron風スケジューラ:1分ごとの異常監視、6時間ごとのメモリ構築 プロバイダ非依存 —...
xalgorix
Xalgorix — Open-source AI pentester that proves vulnerabilities Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported — so you get proof, not a pile of maybes to triage...
autonomous-offensive-llm-handbook
📖 Read the handbook online : follow the lessons from your first fixture run to an agent with clear limits. The model proposes, the code disposes AUTONOMOUS SECURITY / THE BUILDER'S FIELD MANUAL Read the website · Start the course · Run the lab · Connect your model · The build sequence A practical...
seclab-taskflows-fuzzing
Seclab Taskflows Fuzzing An LLM-driven, OSS-Fuzz-style fuzzing pipeline for native C/C++ projects. AFL++ for execution, clang+lcov for coverage, an LLM agent for harness writing, coverage-feedback decisions, triage, and reporting. Fully autonomous: give it a GitHub repo and it handles everything...
Xalgorix Autonomous AI Pentesting Agent 4.6.136
Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported - so you get proof, not a pile of maybes to triage. Self-hosted, private, and bring-your-own-LLM. Built in Go + TypeScript...
llm-agent-testbed
🛡️ LLM Agent Security Testbed Empirical Vulnerability & Defense Harness for Tool-Calling LLM Agents A disciplined security testbed testing whether tool-equipped LLM agents can be manipulated into unauthorized data exfiltration via prompt injection, role-claim social engineering, and confused-deput...
prompt_injection
Cross-Surface Prompt Injection Benchmark Benchmark for measuring where in a tool-using LLM agent's execution pipeline a prompt injection defense fires — and where it doesn't. Embeds unique canary tokens SECRET-A-F0-98 in injected payloads and tracks them at four pipeline stages: exposed → persist...
ABSENTIA: Detecting Broken Access Control Vulnerabilities in Web Applications
Broken access control, the failure of authorization, is one of the most prevalent web security risks. Unlike injection, a flow of untrusted input into a dangerous operation, authorization is a relation: who may act on what, not how data moves. Each application decides that relation for itself, so...
T-MAP
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search T-MAP is a trajectory-aware evolutionary search framework for red-teaming LLM agents over MCP servers. It iteratively generates and mutates adversarial prompts guided by execution trajectories, mapping the agent's vulnerabili...