2 matches found
BadReward: Clean-Label Poisoning of Reward Models in Text-To-Image RLHF
Reinforcement Learning from Human Feedback RLHF is crucial for aligning text-to-image T2I models with human preferences. However, RLHF's feedback mechanism also opens new pathways for adversaries. This paper demonstrates the feasibility of hijacking T2I models by poisoning a small fraction of...
6.9AI score
SaveExploits0
MAI-2024-0017
Large Language Models LLMs utilizing Reinforcement Learning from Human Feedback RLHF for safety alignment are susceptible to a sophisticated "alignment-based" jailbreak attack. This attack employs a best-of-N sampling strategy in conjunction with an adversarial LLM to efficiently craft prompts th...
8.7CVSS5.8AI score
SaveExploits0References1
20