1 matches found
Mechanistic Interpretability of LLM Jailbreaks Via Internal Attribution Graphs
Large language models LLMs exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution methods, offering limited insight into how adversarial...
6.1AI score
SaveExploits0
20