Lucene search
+L

1 matches found

Mend
Mend
•added 2024/12/01 10:00 p.m.•1 views

MAI-2024-0017

Large Language Models LLMs utilizing Reinforcement Learning from Human Feedback RLHF for safety alignment are susceptible to a sophisticated "alignment-based" jailbreak attack. This attack employs a best-of-N sampling strategy in conjunction with an adversarial LLM to efficiently craft prompts th...

8.7CVSS
SaveExploits0References1
Rows per page
Query Builder