3 matches found
AutoRAN-public
🧠 AutoRAN: Secuestro automatizado del razonamiento de seguridad en grandes modelos de razonamiento AutoRAN es un secuestro automatizado del razonamiento de seguridad que aprovecha modelos auxiliares secundarios menos alineados para simular trazas de razonamiento, generar prompts narrativos y...
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can gradually steer the conversation toward an unsafe answer. Recent methods demonstrate this risk, but they are usually...
AutoRAN: Weak-To-Strong Jailbreaking of Large Reasoning Models
This paper presents AutoRAN, the first automated, weak-to-strong jailbreak attack framework targeting large reasoning models LRMs. At its core, AutoRAN leverages a weak, less-aligned reasoning model to simulate the target model's high-level reasoning structures, generates narrative prompts, and...