1 matches found
Red-Teaming Auto Mode: Improving Blocking Classifiers against Malign Coding Agents
To keep coding agents from going off the rails, production systems now review each proposed action with a blocking monitor that can reject it before it runs Auto Mode in Claude Code, Guardian in OpenAI's Codex. Prior evaluations of such monitors largely measure robustness to accidental harm or...
5.9AI score
SaveExploits0
20