1 matches found
Jailbreaking Open-Weight LLMs Via Random Embedding Perturbations
While open-weight models have enjoyed steady progress in capabilities and wide adoption across multiple domains, their safety remains an important concern. One key feature is the ability to refuse or deflect harmful, malicious, or insensitive prompts. In this paper, we expose safety vulnerabiliti...
5.8AI score
SaveExploits0
20