1 matches found
AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation
Safety evaluation of large language models LLMs relies largely on single-turn attack datasets and single-judge scoring, underestimating risk from adaptive multi-turn adversaries and reporting a single success rate that does not separate partially actionable outputs from those carrying complete...
6.1AI score
SaveExploits0
20