1 matches found
Backdoor Containment Via Expert Quarantine and Shutdown in LLMs
Backdoored large language models LLMs can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either...
5.8AI score
SaveExploits0
20