Lucene search
+L

1 matches found

Packet Storm News
Packet Storm News
added 2025/05/17 12:0 a.m.8 views

Self-Destructive Language Model

Harmful fine-tuning attacks pose a major threat to the security of large language models LLMs, allowing adversaries to compromise safety guardrails with minimal harmful data. While existing defenses attempt to reinforce LLM alignment, they fail to address models' inherent "trainability" on harmfu...

7.3AI score
SaveExploits0
Rows per page
Query Builder