Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2025/10/16 12:0 a.m.8 views

PoTS: Proof-Of-Training-Steps for Backdoor Detection in Large Language Models

As Large Language Models LLMs gain traction across critical domains, ensuring secure and trustworthy training processes has become a major concern. Backdoor attacks, where malicious actors inject hidden triggers into training data, are particularly insidious and difficult to detect. Existing...

7.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/22 12:0 a.m.6 views

Unlearning Isn'T Deletion: Investigating Reversibility of Machine Unlearning in LLMs

Unlearning in large language models LLMs is intended to remove the influence of specific data, yet current evaluations rely heavily on token-level metrics such as accuracy and perplexity. We show that these metrics can be misleading: models often appear to forget, but their original behavior can ...

6.6AI score
SaveExploits0
Rows per page
Query Builder