1 matches found
A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks
Large language models LLMs remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs. As jailbreak strategies keep emerging, defenses have proliferated in an ongoing cat-and-mouse gam...
5.5AI score
SaveExploits0
20