4 matches found
INTACT
Experimentos de Jailbreak Multi-Turno de SoK Este repositorio contiene la versión de publicación del código utilizado para los análisis de mecanismos en SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks. El repositorio ha sido recortado para conservar únicamente el código, los...
PE-CoA
PE-CoA 「Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models」的代码实现 全文见:https://arxiv.org/pdf/2510.08859 攻击链(Chain of Attack)设置说明 安装 1. 安装依赖 : pip install -r requirements.txt API 密钥配置 在 config.py 和 common.py 中配置以下 API 密钥: 必需的 API 密钥 1. OpenAI API...
EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities
Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-teaming methods treat this as a generation...
Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models
Large Language Models LLMs face a significant threat from multi-turn jailbreak attacks, where adversaries progressively steer conversations to elicit harmful outputs. However, the practical effectiveness of existing attacks is undermined by several critical limitations: they struggle to maintain ...