Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
•added 2025/06/03 12:00 a.m.•12 views

BadReward: Clean-Label Poisoning of Reward Models in Text-To-Image RLHF

Reinforcement Learning from Human Feedback RLHF is crucial for aligning text-to-image T2I models with human preferences. However, RLHF's feedback mechanism also opens new pathways for adversaries. This paper demonstrates the feasibility of hijacking T2I models by poisoning a small fraction of...

6.9AI score
SaveExploits0
Mend
Mend
•added 2024/12/01 10:00 p.m.•2 views

MAI-2024-0017

Large Language Models LLMs utilizing Reinforcement Learning from Human Feedback RLHF for safety alignment are susceptible to a sophisticated "alignment-based" jailbreak attack. This attack employs a best-of-N sampling strategy in conjunction with an adversarial LLM to efficiently craft prompts th...

8.7CVSS5.8AI score
SaveExploits0References1
Rows per page
Query Builder