1 matches found
Circuit Discovery Helps Detect LLM Jailbreaking: A Mechanistic Interpretability Study
Despite extensive safety alignment, large language models LLMs remain vulnerable to jailbreak attacks that bypass safeguards to elicit harmful content. While prior work attributes this vulnerability to safety training limitations, the internal mechanisms by which LLMs process adversarial prompts...
5.8AI score
SaveExploits0
20