Lucene search
+L

1 matches found

Packet Storm News
Packet Storm News
added 2025/06/19 12:0 a.m.7 views

Probe Before You Talk: Towards Black-Box Defense against Backdoor Unalignment for Large Language Models

Backdoor unalignment attacks against Large Language Models LLMs enable the stealthy compromise of safety alignment using a hidden trigger while evading normal safety auditing. These attacks pose significant threats to the applications of LLMs in the real-world Large Language Model as a Service...

7.4AI score
SaveExploits0
Rows per page
Query Builder