2 matches found
Which Neurons Detect Malicious Code? A Probing Study of LLM Security Knowledge
Background. Large language models LLMs have become increasingly capable of understanding and generating source code, leading to their widespread adoption in software engineering tasks such as code completion, repair, and vulnerability detection. However, despite their strong empirical performance...
Mechanistic Interpretability of LLM Jailbreaks Via Internal Attribution Graphs
Large language models LLMs exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution methods, offering limited insight into how adversarial...