1 matches found
MAI-2024-0017
Large Language Models LLMs utilizing Reinforcement Learning from Human Feedback RLHF for safety alignment are susceptible to a sophisticated "alignment-based" jailbreak attack. This attack employs a best-of-N sampling strategy in conjunction with an adversarial LLM to efficiently craft prompts th...
8.7CVSS
SaveExploits0References1
20