2 matches found
Jailbreaking Open-Weight LLMs Via Random Embedding Perturbations
While open-weight models have enjoyed steady progress in capabilities and wide adoption across multiple domains, their safety remains an important concern. One key feature is the ability to refuse or deflect harmful, malicious, or insensitive prompts. In this paper, we expose safety vulnerabiliti...
5.8AI score
SaveExploits0
Fine-Grained Distributed Backdoor Attacks in Federated Learning
Federated learning, as a privacy-preserving distributed machine learning paradigm, faces significant threats from backdoor attacks. Compared to centralized attacks, distributed backdoor attacks are more harmful but require more poisoned samples to compensate for the loss of trigger strength due t...
5.9AI score
SaveExploits0
20