1 matches found
Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs Via Weight Orthogonalisation
Backdoor attacks can be implanted in Large Language Models LLMs during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can...
5.7AI score
SaveExploits0
20