3 matches found
AutoRAN-public
🧠 AutoRAN: Secuestro automatizado del razonamiento de seguridad en grandes modelos de razonamiento AutoRAN es un secuestro automatizado del razonamiento de seguridad que aprovecha modelos auxiliares secundarios menos alineados para simular trazas de razonamiento, generar prompts narrativos y...
SHIELD: a Secure and Highly Enhanced Integrated Learning for Robust Deepfake Detection against Adversarial Attacks
Audio plays a crucial role in applications like speaker verification, voice-enabled smart devices, and audio conferencing. However, audio manipulations, such as deepfakes, pose significant risks by enabling the spread of misinformation. Our empirical analysis reveals that existing methods for...
Prediction Inconsistency Helps Achieve Generalizable Detection of Adversarial Examples
Adversarial detection protects models from adversarial attacks by refusing suspicious test samples. However, current detection methods often suffer from weak generalization: their effectiveness tends to degrade significantly when applied to adversarially trained models rather than naturally train...