Lucene search
+L

2 matches found

Kitploit
Kitploit
added 2026/09/10 7:56 p.m.6 views

AutoRAN-public

🧠 AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models AutoRAN is an automated Hijacking of Safety Reasoning that leverages less-aligned secondary auxiliary models to simulate reasoning traces, generate narrative prompts, and iteratively refine those prompts to bypass safety...

6AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/21 12:00 a.m.19 views

SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning

Large Reasoning Models LRMs introduce a new generation paradigm of explicitly reasoning before answering, leading to remarkable improvements in complex tasks. However, they pose great safety risks against harmful queries and adversarial attacks. While recent mainstream safety efforts on LRMs,...

7.3AI score
SaveExploits0
Rows per page
Query Builder