3 matches found
StealthRL
StealthRL: Ataques de Paráfrasis mediante Aprendizaje por Refuerzo para la Evasión Multi-Detector de Detectores de Texto IA Artículo arXiv Demo Modelo Hugging Face Dataset de Referencia Hugging Face Resumen Los detectores de texto IA se utilizan cada vez más en entornos de alto riesgo, pero su...
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
AI-text detectors face a critical robustness challenge: adversarial paraphrasing attacks that preserve semantics while evading detection. We introduce StealthRL, a reinforcement learning framework that stress-tests detector robustness under realistic adversarial conditions. StealthRL trains a...
GradEscape: a Gradient-Based Evader against AI-Generated Text Detectors
In this paper, we introduce GradEscape, the first gradient-based evader designed to attack AI-generated text AIGT detectors. GradEscape overcomes the undifferentiable computation problem, caused by the discrete nature of text, by introducing a novel approach to construct weighted embeddings for t...