1 matches found
EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities
Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-teaming methods treat this as a generation...
5.9AI score
SaveExploits0
20