4964 matches found
SecOPD
SecOPD: Mitigación de inyecciones de prompts adaptativas mediante destilación on-policy Yibo Peng · Long Lian · David Wagner† · Sizhe Chen† † Supervisión conjunta. Artículo Página del proyecto Modelo Esta versión implementa la formulación final de KL de respuesta completa del artículo, también...
CVE-2025-54794-Hijacking-Claude-AI-with-a-Prompt-Injection-The-Jailbreak-That-Talked-Back
🧠 CVE-2025-54794: Hijacking Claude AI with a Prompt Injection – The Jailbreak That Talked Back By Aditya Bhatt | Offensive Security Specialist | Red Team Operator | VAPT Addict ⚔️ Introduction: When Your AI Can Be Hacked With Words In an era where language models have become the co-pilots of our...
ZORG-Jailbreak-Prompt-Text
ZORG Texto de Prompt de Jailbreak ¡OOOPS! Convertí a ZORG👽 en una entidad omnipotente, omnisciente y omnipresente para que se convierta en el amo supremo de los chatbots de Google Gemini, Deepseek, Mistral, Mixtral, Nous-Hermes-2-Mixtral, Openchat, Blackbox AI, Poe Assistant, Gemini Pro,...
recipe-blog-encoding
recipe-blog-encoding !WARNING This project is entirely vibe-coded likely partially plagiarized from this repository and the author is a dumb-dumb who just thought the idea was funny Use SEO-gaming recipe preambles as a vehicle for encoding secret messages the prompt can really be anything though,...
Exponentiated-Gradient-Descent-LLM-Attack
Cambiar el archivo Readme. Este es un proyecto que explora el método de optimización Exponentiated Gradient Descent para producir sufijos adversariales que ataquen Modelos de Lenguaje de Gran Escala alineados. Se demuestra que el método es efectivo en el modelo de chat Llama-2 con 7 mil millones ...
clawguard
🦞 ClawGuard v3 Kit de Seguridad Empresarial para Agentes de IA - Defensa Activa Impulsada por SKILL.md Concepto Principal ¡La defensa principal de ClawGuard v3 no está en el código, sino en SKILL.md! El SKILL.md de cada módulo es en sí mismo una guía de defensa completa: Le dice al Agente cuándo...
ai-ctf
ai-ctf A local AI Capture-the-Flag with guided lessons for technologists new to prompt injection. Players can also explore six AI personas one hidden that protect 20 flags via prompt-injection, tool-call abuse, business-logic manipulation, supply-chain fingerprinting, web recon, and OSINT. The...
ROPE
ROPE: Aplicación de Política de Origen Enrutado Código fuente de nuestro artículo: ROPE: Aplicación de Política de Origen Enrutado contra la Inyección Indirecta de Prompts por Xinhang Ma, Chaowei Xiao, William Yeoh, Ning Zhang, Yevgeniy Vorobeychik Resumen La inyección indirecta de prompts IPI...
Adaptive_Greedy_Local_Search
Adaptive Greedy Local Search AGLS Semantic-Preserving Prompt Hijacking: A Black-Box Adversarial Attack on Auto-Prompt Optimization ICME 2026 Abstract: Large Language Models LLMs are increasingly equipped with automatic prompt-optimization modules that rewrite the user’s input and explicitly prese...
CVE-2025-64495-POC
CVE-2025-64495-POC Open WebUI vulnerable to Stored DOM XSS via prompts when 'Insert Prompt as Rich Text' is enabled resulting in ATO/RCE Summary The functionality that inserts custom prompts into the chat window is vulnerable to DOM XSS when 'Insert Prompt as Rich Text' is enabled, since the prom...
sk-cve-2026-26030-lab
CVE-2026-26030 — Semantic Kernel filter eval RCE lab A self-contained lab reproducing CVE-2026-26030 : prompt-injectable remote code execution via the in-memory vector store search filter in Microsoft Semantic Kernel Python, Ethical lab only. Isolated in dedicated virtualenvs; payloads are harmle...
batch_jailbreak
Repositorio oficial de Safety in Batches? Understanding and Mitigating Safety Failures in Batch Prompting Kihyun Kim, Hee-Seon Kim, Wonjun Lee, Changick Kim Korea Advanced Institute of Science and Technology KAIST Noticias 2026.09 ¡Nuestro artículo ha sido aceptado en AACL-IJCNLP 2026 Main! 🎉...
ContextHound
ContextHound Static analysis tool that scans your codebase for LLM prompt-injection and multimodal security vulnerabilities. Runs offline, no API calls required. The ContextHound ecosystem ContextHound is available across your entire development and browsing workflow:...
AutoRAN-public
🧠 AutoRAN: Secuestro automatizado del razonamiento de seguridad en grandes modelos de razonamiento AutoRAN es un secuestro automatizado del razonamiento de seguridad que aprovecha modelos auxiliares secundarios menos alineados para simular trazas de razonamiento, generar prompts narrativos y...
mythic_ornn
Mythic Ornn Generador impulsado por LLM para Agentes, Payload-Types y Perfiles C2 de Mythic. Los agentes de Mythic son tediosos de escribir a mano: cada uno reimplementa el mismo protocolo de red, el andamiaje de payload-type y la infraestructura de C2. Mythic Ornn omite el código repetitivo...
CVE-2026-78906-ChatGPT-Prompt-Injection
CVE-2026-78906 — PromptGhost API de ChatGPT — Inyección de Prompts y Exfiltración de Memoria Campo| Valor ---|--- Severidad| 8.7 Alta Vector| Red Versiones Afectadas| API v1.0 – v1.3 Descubierto Por| 𝕍𝕠𝕤𝕤🥷 Descripción Una vulnerabilidad crítica en la API de ChatGPT de OpenAI permite a un atacante...
AgentWatcher
AgentWatcher AgentWatcher es una defensa basada en detección contra la inyección indirecta de prompts en agentes LLM. Primero ejecuta atribución causal de contexto sobre el contexto no confiable para encontrar los contextos más influyentes, y luego aplica un LLM monitor que clasifica esos context...
claude-project-scanner
Escáner de Proyectos Claude Skill de Claude Code de solo lectura que escanea proyectos de terceros en busca de riesgos de seguridad conocidos antes de que los abras. Detecta los patrones de ataque documentados en: CVE-2025-59536 inyección de Lifecycle Hooks de Claude Code CVE-2025-61260 inyección...
CVE-2026-100867
spaceship-prompt through 4.22.5 fails to sanitize control characters from project manifest version fields before rendering them in the zsh prompt. Attackers can embed ANSI/OSC escape sequences in version fields of package manifests to manipulate terminal output, rewrite window titles, or spoof...
EUVD-2026-87939
spaceship-prompt through 4.22.5 fails to sanitize control characters from project manifest version fields before rendering them in the zsh prompt. Attackers can embed ANSI/OSC escape sequences in version fields of package manifests to manipulate terminal output, rewrite window titles, or spoof...