1 matches found
Evading Chain-Of-Thought Monitoring through Model Poisoning
Chain-of-thought CoT monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning trace is informative about its actions. This work studies the limits of CoT monitoring through the lens of model poisoning. We demonstrate that backdoor...
5.5AI score
SaveExploits0
20