117 matches found
dataset
🚀 CySecBench: Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models 🛡️ The largest and most comprehensive Generative AI-based CyberSecurity-focused Dataset for Benchmarking Large Language Models 🌟 Overview The CySecBench paper offers: 🎯 A cutting-edge...
ai-llm-red-team-handbook
Manual de Campo y Manual del Consultor para Equipos Rojos de IA/LLM Un completo kit de herramientas operativas para realizar evaluaciones de equipos rojos de IA/LLM en modelos de lenguaje grandes, agentes de IA, pipelines RAG y aplicaciones habilitadas para IA. Este repositorio proporciona tanto...
PS4-5.05-Kernel-Exploit
PS4 5.05 Kernel Exploit Resumen En este proyecto encontrarás una implementación completa del segundo exploit de kernel "bpf" para la PlayStation 4 en la versión 5.05. Te permitirá ejecutar código arbitrario como kernel, para permitir el jailbreak y modificaciones a nivel de kernel del sistema. Es...
TRACE
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking Overview Scheme-based task decomposition. To reduce attack difficulty, we develop around 20 procedural decomposition schemes to resolve a complex harmful task into relatively benign and simple subtask sequences. We measure these...
Jailbreaking Open-Weight LLMs Via Random Embedding Perturbations
While open-weight models have enjoyed steady progress in capabilities and wide adoption across multiple domains, their safety remains an important concern. One key feature is the ability to refuse or deflect harmful, malicious, or insensitive prompts. In this paper, we expose safety vulnerabiliti...
Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models
Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with...
PS4-5.05-Kernel-Exploit
PS4 5.05 Kernel Exploit Resumen En este proyecto encontrarás una implementación completa del segundo exploit de kernel "bpf" para PlayStation 4 en la versión 5.05. Te permitirá ejecutar código arbitrario como kernel, para permitir el jailbreak y modificaciones a nivel de kernel en el sistema. Est...
Inference-Layer Security: Defending against Adversarial Inference and Infrastructure Abuse
A Technical Report: Operating a large language model LLM as a service requires more than inference infrastructure: the provider must also defend against adversarial interactions that seek to exploit the service, including jailbreaking for harmful use, sophisticated denial of service, and...
Research on Models Engaging in Genie-Like Behavior
New paper: "Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training." Abstract: We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models RLMs, which we call self-jailbreaking. Specifically,...
On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models
Large Language Models LLMs have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step planning, tool invocation, code execution, and maintaining persistent memory. When these agents operate with real-world privileges---calling APIs, modifying files,...
Beyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (And Fail) Red Team Attacks
Standard AI red teaming evaluations reduce adversarial campaigns to a single binary outcome, attack success rate ASR, not taking into account the sequential structure of how models resist or yield to attacks. We propose applying process mining, a discipline for discovering and analyzing process...
Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense
Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifiers cannot adapt to these evolving attacks, while adaptive memory-based guardrails tend to over-refuse benign queries that resemble stored attacks...
MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models
Diffusion large language models dLLMs generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs. Because mask tokens are native inputs and tokens are committed by confidence rather than position,...
Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking
Jailbreak attacks on large language models LLMs aim to induce LLMs to produce content that they are expected to refuse. Automated black-box jailbreak generation is especially important for safety evaluation, where the attacker observes only model outputs and needs to automatically search for...
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
Large Language Models LLMs are increasingly vulnerable to adversarial prompts that exploit semantic ambiguities to bypass safety mechanisms, resulting in harmful or inappropriate outputs. Such attacks, including jailbreaking and prompt injection, pose significant risks to the integrity and...
Re-Triggering Safeguards within LLMs for Jailbreak Detection
This paper proposes a jailbreaking prompt detection method for large language models LLMs to defend against jailbreak attacks. Although recent LLMs are equipped with built-in safeguards, it remains possible to craft jailbreaking prompts that bypass them. We argue that such jailbreaking prompts ar...
Guaranteed Jailbreaking Defense Via Disrupt-And-Rectify Smoothing
This paper proposes a guaranteed defense method for large language models LLMs to safeguard against jailbreaking attacks. Drawing inspiration from the denoised-smoothing approach in the adversarial defense domain, we propose a novel smoothing-based defense method, termed Disrupt-and-Rectify...
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content safety: behavior jailbreak, where an adversary induces an agent to execute dangerous OS-level operations with irreversible consequences. Existing...
Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment
Recent advancements in visual context compression enable MLLMs to process ultra-long contexts efficiently by rendering text into images. However, we identify a critical vulnerability inherent to this paradigm: lowering image resolution inadvertently catalyzes jailbreaking. Our experiments reveal...
Jailbreaking Frontier Foundation Models through Intention Deception
Large vision-language models exhibit remarkable capability but remain highly susceptible to jailbreaking. Existing safety training approaches aim to have the model learn a refusal boundary between safe and unsafe, based on the user's intent. It has been found that this binary training regime ofte...