Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2026/05/21 12:0 a.m.17 views

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications

Large Language Models LLMs have become the predominant paradigm in NLP, advancing both research and industry. As model sizes and pretraining data grow, concerns about Pretraining Data Exposure PDE increase due to the scale and opacity of training datasets. PDE refers to determining whether specif...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/08 12:0 a.m.5 views

STAMP Your Content: Proving Dataset Membership Via Watermarked Rephrasings

Given how large parts of publicly available text are crawled to pretrain large language models LLMs, data creators increasingly worry about the inclusion of their proprietary data for model training without attribution or licensing. Their concerns are also shared by benchmark curators whose...

6.9AI score
SaveExploits0
Rows per page
Query Builder