1 matches found
AEGIS: Audio Endogenous Guarding Via Internal Signals against Large Audio-Language Model Jailbreaks
Large audio-language models LALMs expand language models to process and interpret audio, but also expose them to heterogeneous audio jailbreaks. We ask whether successful jailbreaks reflect failures to recognize harmful intent or failures occurring after such recognition. Layer-wise probing revea...
5.8AI score
SaveExploits0
20