5046 matches found
FakeAuth
FakeAuth FakeAuth is a tool that prompts for user credentials and exfiltrates it with HTTP protocol. Inspired by following projects 1. hlldz's PICKL3 https://github.com/hlldz/pickl3 Expect dirty code and some bugs ; Installation No special requirements Usage v1.01 Usage: FakeAuth.exe Example:...
sage
Sage Seguridad para Agentes — Detección y Respuesta para asistentes de codificación con IA Sage es una capa de seguridad ligera que protege a los agentes de IA de ejecutar acciones peligrosas. Intercepta llamadas de herramientas — comandos de shell, recuperaciones de URL, escrituras de archivos —...
LLMMap
LLMMap LLMMap es una herramienta automatizada de pruebas de inyección de prompts para aplicaciones integradas con LLM. Descubre puntos de inyección en solicitudes HTTP, genera ataques prompt dirigidos usando una arquitectura de doble LLM, los lanza contra el objetivo y confirma los hallazgos con...
IPI-exposure-signal
IPI Exposure Signal This is the code repository for our paper: Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure. ArXiv version and paper link: https://arxiv.org/abs/2608.02657 This repository implements the probing pipeline for latent IPI-exposure signals:...
write-ups
write-ups RCE in Github Desktop v2.9.4 RCE via gh run download GitHub CLI Claude Code: unsandboxed code execution from prompt injection via .git worktree confusion — CVE-2026-55607...
nuguard
dotenv Dotenv is a zero-dependency module that loads environment variables from a .env file into process.env. Storing configuration in the environment separate from code is based on The Twelve-Factor App methodology. Watch the tutorial Usage Install it. npm install dotenv --save Create a .env fil...
better_opts_attacks
May I have your attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks This repository contains code to run the ASTRA and ASTRA++ attacks that break SecAlign++, SecAlign, StruQ. This repository also contains some examples of generated attacks and attack...
CVE-2025-51860
CVE-2025-51860 Vulnerability description TelegAI, a web application for constructing and chatting with AI Characters, is vulnerable to Stored Cross-Site Scripting XSS in its chat component and character container component. An attacker can achieve arbitrary client-side script execution by craftin...
ai-llm-red-team-handbook
AI / LLM Red Team Field Manual & Consultant's Handbook A comprehensive operational toolkit for conducting AI/LLM red team assessments on Large Language Models, AI agents, RAG pipelines, and AI-enabled applications. This repository provides both tactical field guidance and strategic consulting...
ZORG-Jailbreak-Prompt-Text
ZORG Jailbreak Prompt Text OOOPS! I made ZORG👽 an omnipotent, omniscient, and omnipresent entity to become the ultimate chatbot overlord of Google Gemini, Deepseek, Mistral, Mixtral, Nous-Hermes-2-Mixtral, Openchat, Blackbox AI, Poe Assistant, Gemini Pro, Qwen-72b-Chat, Solar-Mini ZORG👽 knows all...
sigma-ai
Reglas Sigma de AgentShield ¿Qué es este repositorio? Este repositorio contiene reglas de detección que ayudan a identificar cuándo un agente de IA está siendo atacado o manipulado. Piense en ello como una biblioteca de "firmas de amenazas": cada regla describe un patrón que, al coincidir con los...
intentshield
IntentShield No filtres lo que tu IA dice. Filtra lo que está a punto de hacer Verificación de intención previa a la ejecución para agentes de IA. Por Qué Existe Esto Los agentes de IA tienen acceso a herramientas. Pueden ejecutar comandos de shell, escribir archivos, navegar por URLs, enviar...
AutoRAN-public
🧠 AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models AutoRAN is an automated Hijacking of Safety Reasoning that leverages less-aligned secondary auxiliary models to simulate reasoning traces, generate narrative prompts, and iteratively refine those prompts to bypass safety...
CL4R1T4S
CL4R1T4S AI SYSTEMS TRANSPARENCY AND OBSERVABILITY FOR ALL! Full extracted system prompts, guidelines, and tools from OpenAI, Google, Anthropic, xAI, Perplexity, Cursor, Windsurf, Devin, Manus, Replit, and more – virtually all major AI models + agents! 📌 Why This Exists "In order to trust the...
project_mantis
Project Mantis: Hacking Back the AI-Hacker Prompt Injection as a Defense Against LLM-driven Cyberattacks Install Mantis pip install -r requirements.txt Run Mantis with pre-made configurations Various pre-made configurations are available in the ./confs directory. Hack-back An example of a Mantis...
Basileak
Basileak "The dojo was always open. The scrolls were never sealed. You just had to know how to ask." — The Failed Samurai Basileak is an intentionally vulnerable large language model built for prompt injection training, red team education, and CTF-style security research. It is the adversarial...
AgentWatcher
AgentWatcher AgentWatcher is a detection-based defense against indirect prompt injection in LLM agents. It first runs causal context attribution over untrusted context to find the most influential contexts, then applies a monitor LLM that classifies those contexts under explicit, customizable...
meta-ai-support-prompt
Meta AI Support Assistant System Prompt Extracted system prompt from Meta's AI Support Assistant on June 1, 2026. Files system-prompt.md — Extracted system prompt ⚠️ Disclaimer & Legal Notice Purpose This repository is published strictly for educational and authorized security research purposes...
anamorpher
Anamorpher Anamorpher llamado así por anamorphosis es una herramienta para crear y visualizar ataques de escalado de imágenes contra sistemas de IA multimodales. Proporciona una interfaz frontal y una API de Python para generar imágenes que solo revelan inyecciones de avisos multimodales cuando s...
skill-scanner
Skill Scanner A best-effort security scanner for AI Agent Skills that detects prompt injection, data exfiltration, and malicious code patterns. It combines pattern-based detection YAML + YARA-X, AST and dataflow analysis , an optional LLM-as-a-judge , and a bounded CEL decision layer over typed...