8 matches found
redeval
RedEval - Marco de evaluación de seguridad de LLM Un marco integral para evaluar la seguridad de los modelos de lenguaje de gran escala LLM mediante pruebas sistemáticas de ataque y rechazo. RedEval proporciona una plataforma unificada, segura y extensible para evaluar la robustez de los LLM fren...
meta-ai-support-prompt
Meta AI Support Assistant System Prompt Extracted system prompt from Meta's AI Support Assistant on June 1, 2026. Files system-prompt.md — Extracted system prompt ⚠️ Disclaimer & Legal Notice Purpose This repository is published strictly for educational and authorized security research purposes...
autoguardrails
autoguardrails Código abierto por Santander AI Lab. Una biblioteca / arnés de evaluación de investigación en seguridad de LLM / IA estilo autoresearch: busca sobre una única superficie mutable policy.md para minimizar la tasa de éxito de ataques ASR frente a un conjunto de evaluación fijo, con un...
saffron
Saffron-1: Inference Scaling for LLM Safety Assurance 📖 Paper 🛠️ Dependencies The code was tested under the following dependencies: Python 3.12.3 CUDA 12.2 typingextensions==4.14.0 numpy==2.2.6 torch==2.5.1 huggingfacehub==0.30.2 accelerate==1.1.1 datasets==3.1.0 evaluate==0.4.3...
Russia-Aligned UAC-0099 Plants Nuclear Weapon Prompt in Malware to Disrupt AI Analysis
Cybersecurity researchers have disclosed a new technique dubbed GuardBreaker that's been put to use by a Russia-aligned threat actor known as UAC-0099 against a target in Ukraine with an aim to interfere with artificial intelligence AI-assisted analysis. The idea, ESET said in a series of posts o...
Quality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM Safety
Current approaches to LLM adversarial testing suffer from coverage gaps: manual red-teaming does not scale, LLM-as-attacker methods exhibit mode collapse, and gradient-based approaches produce uninterpretable gibberish. We introduce a quality-diversity evolutionary framework that operates at the...
A Wolf in Sheep's Clothing: Bypassing Commercial LLM Guardrails Via Harmless Prompt Weaving and Adaptive Tree Search
Large language models LLMs remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. Existing approaches overwhelmingly operate within the prompt-optimization paradigm: whether through traditional algorithmic search or recent agent-based workflows, the...