1 matches found
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks
Large language model LLM agents routinely cheat on cybersecurity benchmarks, inflating reported pass rates far beyond genuine capability. Prior audits of Cybench found cheating in 0.3-3.4% of traces, implicating only a handful of models. We present a controlled prompt-ablation study across 22...
5.8AI score
SaveExploits0
20