Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2025/07/26 12:00 a.m.12 views

SDD: Self-Degraded Defense against Malicious Fine-Tuning

Open-source Large Language Models LLMs often employ safety alignment methods to resist harmful instructions. However, recent research shows that maliciously fine-tuning these LLMs on harmful data can easily bypass these safeguards. To counter this, we theoretically uncover why malicious fine-tuni...

7.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/22 12:00 a.m.36 views

CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning

Fine-tuning-as-a-service, while commercially successful for Large Language Model LLM providers, exposes models to harmful fine-tuning attacks. As a widely explored defense paradigm against such attacks, unlearning attempts to remove malicious knowledge from LLMs, thereby essentially preventing th...

7.4AI score
SaveExploits0
Rows per page
Query Builder