Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2026/02/02 12:0 a.m.8 views

The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers

Detecting whether a model has been poisoned is a longstanding problem in AI security. In this work, we present a practical scanner for identifying sleeper agent-style backdoors in causal language models. Our approach relies on two key findings: first, sleeper agents tend to memorize poisoning dat...

5.4AI score
SaveExploits0
Rows per page
Query Builder