Lucene search
+L

1 matches found

Packet Storm News
Packet Storm News
added 2025/06/03 12:0 a.m.7 views

BadReward: Clean-Label Poisoning of Reward Models in Text-To-Image RLHF

Reinforcement Learning from Human Feedback RLHF is crucial for aligning text-to-image T2I models with human preferences. However, RLHF's feedback mechanism also opens new pathways for adversaries. This paper demonstrates the feasibility of hijacking T2I models by poisoning a small fraction of...

6.9AI score
SaveExploits0
Rows per page
Query Builder