2 matches found
Anthropic’s Opus 5 Is Better at Resisting Prompt Injection
The chart is interesting. On the IPI benchmark, Opus 5 improved over Opus 4.8, reducing the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0%, and from 0.5% to 0.2% on 1 attempt. It also improved on Sonnet 5 5.9% at k=15 and Mythos 5 2.6%, making it the most robust model...
5.5AI score
SaveExploits0
A Red-Team Study of Anthropic Fable 5 and Opus 4.8 Models
We evaluate the adversarial robustness of two frontier large language models LLMs developed by Anthropic, Fable 5 and Opus 4.8, against four families of automated jailbreak attack across 7 826 harmful intents spanning a ten-category harm taxonomy. Using the HackAgent red-teaming framework, hundre...
5.4AI score
SaveExploits0
20