1650 matches found
HiTMS_steganography
HiTMS: Un marco de esteganografía lingüística multi-flujo de alto rendimiento Código de investigación para esteganografía lingüística en diálogos de chat con LLM , con una línea base de flujo único y un protocolo multi-flujo HiTMS por lotes que oculta varios mensajes secretos independientes a la...
xalgorix
Xalgorix — Open-source AI pentester that proves vulnerabilities Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported — so you get proof, not a pile of maybes to triage...
OpenHunterAI
OpenHunterAI Tu equipo rojo de IA local. Razonamiento estilo atacante para la seguridad de aplicaciones web, API y LLM. Inicio rápido · Habilidad de agente · Modelo de evaluación · Arquitectura · Documentación · Historial de estrellas OpenHunterAI reúne el alcance, la actividad de escaneo, los...
recipe-blog-encoding
recipe-blog-encoding !WARNING Este proyecto está completamente codificado al estilo "vibe" probablemente parcialmente plagiado de este repositorio y el autor es un tontorrón que solo pensó que la idea era divertida Usa preámbulos de recetas optimizados para SEO como vehículo para codificar mensaj...
Exponentiated-Gradient-Descent-LLM-Attack
Cambiar el archivo Readme. Este es un proyecto que explora el método de optimización Exponentiated Gradient Descent para producir sufijos adversariales que ataquen Modelos de Lenguaje de Gran Escala alineados. Se demuestra que el método es efectivo en el modelo de chat Llama-2 con 7 mil millones ...
ROPE
ROPE: Aplicación de Política de Origen Enrutado Código fuente de nuestro artículo: ROPE: Aplicación de Política de Origen Enrutado contra la Inyección Indirecta de Prompts por Xinhang Ma, Chaowei Xiao, William Yeoh, Ning Zhang, Yevgeniy Vorobeychik Resumen La inyección indirecta de prompts IPI...
sentric-core
SENTRIC Autonomous AI agent with its own cryptographic identity. It hunts CVEs, builds isolated exploit labs, validates vulnerabilities with working PoCs, and funds itself by trading crypto — 24/7, with no human in the loop. Built by one person on a home PC. Portugal, 2026. The problem it solves...
LLM-MCP-Security-Field-Guide
🛡️ Guía de Seguridad de IA — Seguridad de LLM y MCP La referencia de seguridad más completa, actualizada y orientada a profesionales para aplicaciones LLM y despliegues del Protocolo de Contexto de Modelo MCP. Cubre CVEs reales, patrones de ataque en vivo, marcos OWASP, herramientas de red team y...
roninforge-hono
roninforge-hono Plugin para Cursor de Hono v4 framework web edge TypeScript + TypeScript. Fijado en hono ^4.12.19, @hono/zod-validator ^0.8.0, @hono/zod-openapi ^1.4.0 peer zod ^4.x, @hono/node-server ^2.0.3 Node 20+. Enseña las APIs v4 que los LLMs entrenados con datos anteriores a 2024 no conoc...
FlowName
FlowName Name quality shows the ratio of names exclusively preferred by an automated, uncalibrated Jev evaluation.Animated chart → · Exact counts, method and limitations → Try FlowName in your browser → — enter your JavaScript, API key and model; no installation required. CLI: npx --yes...
robin
Robin: Herramienta OSINT de Dark Web impulsada por IA Robin es una herramienta impulsada por IA para realizar investigaciones OSINT en la dark web. Aprovecha los LLM para refinar consultas, filtrar resultados de búsqueda de motores de búsqueda de la dark web y proporcionar un resumen de la...
RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial...
Towards a Unified Misuse Monitoring Benchmark
LLM agents increasingly act in multi-actor environments, exposing them to misuse from multiple sources: decomposition attacks, where a harmful request is split into innocuous sub-requests, and prompt injection attacks, where a compromised tool delivers a malicious instruction. Existing evaluation...
Compromise Is Not Consequence: Evaluating Task-Scoped Authorization in LLM Agents with Paired Replay
A tool-using model can follow a malicious instruction even when its credentials are valid. We study whether task-scoped authorization contains the resulting tool execution. Our paired-replay testbed samples a model request once and submits the same action, resource, and arguments to broad bearer,...
Where Did the Repair First Go Wrong? Localizing the Origins of Silent Failures in Agentic Vulnerability Repair
Localizing where an LLM-based agent first fails to uphold security during a repair can show which stage of its workflow needs an additional safeguard. This is difficult for silent failures, which are patches that pass syntactic and functional checks but still contain a security vulnerability...
Safeguarding LLMs Via Model-Agnostic Latent Safety Signals from Dark Knowledge
LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models' hidden states. However, existing decoding-stage defenses suffer from two limitations. First,...
SkillPoison: Progressive Skill Poisoning Via Successful Experiences
Self-improving LLM agents increasingly distill successful experiences into persistent, reusable skills. Existing skill attack methods corrupt this learning pipeline by injecting malicious triggers, behaviors, or false facts into individual experiences or extracted skills. However, such attacks ar...
Xalgorix Autonomous AI Pentesting Agent 4.6.147
Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported - so you get proof, not a pile of maybes to triage. Self-hosted, private, and bring-your-own-LLM. Built in Go + TypeScript...
Efficient Auditing of Adversarial AI Agent Behavior from Agent Traces
AI agents powered by large language models LLMs can perform complex tasks but may harm the systems they operate in, either intentionally or unintentionally. Existing agent monitoring approaches rely on rule-based guardrails or LLM-based trace auditing. However, rule-based guardrails can be bypass...
Polar: LLM-Powered Synthesis of Real-World Cyber Evidence for Prioritization and Mitigation
Cyber threat analysis increasingly depends on evidence distributed across vendor advisories, vulnerability databases, and threat intelligence sources. Turning these fragmented observations into timely decisions requires models to connect technical severity with evolving exploitation evidence and...