Lucene search
+L

5 matches found

Packet Storm News
Packet Storm News
added 2025/06/11 12:0 a.m.5 views

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods

In this work, we show that some machine unlearning methods may fail when subjected to straightforward prompt attacks. We systematically evaluate eight unlearning techniques across three model families, and employ output-based, logit-based, and probe analysis to determine to what extent supposedly...

6.9AI score
SaveExploits0
Rows per page
Query Builder