Lucene search
+L

3 matches found

Kitploit
Kitploit
added 2026/09/09 8:21 a.m.4 views

AutoRAN-public

🧠 AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models AutoRAN is an automated Hijacking of Safety Reasoning that leverages less-aligned secondary auxiliary models to simulate reasoning traces, generate narrative prompts, and iteratively refine those prompts to bypass safety...

6AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/09 12:00 a.m.36 views

MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks

Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can gradually steer the conversation toward an unsafe answer. Recent methods demonstrate this risk, but they are usually...

5.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/16 12:00 a.m.11 views

AutoRAN: Weak-To-Strong Jailbreaking of Large Reasoning Models

This paper presents AutoRAN, the first automated, weak-to-strong jailbreak attack framework targeting large reasoning models LRMs. At its core, AutoRAN leverages a weak, less-aligned reasoning model to simulate the target model's high-level reasoning structures, generates narrative prompts, and...

7.6AI score
SaveExploits0
Rows per page
Query Builder