1 matches found
DITTO: A Spoofing Attack Framework on Watermarked LLMs Via Knowledge Distillation
The promise of LLM watermarking rests on a core assumption that a specific watermark proves authorship by a specific model. We demonstrate that this assumption is dangerously flawed. We introduce the threat of watermark spoofing, a sophisticated attack that allows a malicious model to generate te...
7AI score
SaveExploits0
20