1 matches found
A Red-Team Study of Anthropic Fable 5 and Opus 4.8 Models
We evaluate the adversarial robustness of two frontier large language models LLMs developed by Anthropic, Fable 5 and Opus 4.8, against four families of automated jailbreak attack across 7 826 harmful intents spanning a ten-category harm taxonomy. Using the HackAgent red-teaming framework, hundre...
5.4AI score
SaveExploits0
20