4 matches found
High-Quality Data Do Not Mean Safe! Poisoning LLMs after Data Selection
Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples. However, prior studies typically assume that poisoned samples directly enter downstream fine-tuning, overlooking quality-based selection in practical training pipelines. To fi...
5.9AI score
SaveExploits0
20