2 matches found
aegis-audio-defense
AEGIS: Audio Endogenous Guarding via Internal Signals Against Large Audio-Language Model Jailbreaks Submitted to ICASSP 2027. AEGIS keeps the target audio-language model frozen and adds a mid-layer risk gate that conditionally drives late-layer safety adapters, with closed-loop scaling at...
6.4AI score
SaveExploits0References6
AEGIS: Audio Endogenous Guarding Via Internal Signals against Large Audio-Language Model Jailbreaks
Large audio-language models LALMs expand language models to process and interpret audio, but also expose them to heterogeneous audio jailbreaks. We ask whether successful jailbreaks reflect failures to recognize harmful intent or failures occurring after such recognition. Layer-wise probing revea...
5.8AI score
SaveExploits0
20