2 matches found
StealthRL
StealthRL:用于多检测器规避 AI 文本检测器的强化学习改述攻击 论文(arXiv) 演示 模型(Hugging Face) 基准数据集(Hugging Face) 摘要 AI 文本检测器越来越多地用于高风险场景,但它们对保持语义的对抗性改写的鲁棒性仍不确定。我们提出了 StealthRL,一个通过自适应改述攻击对 AI 文本检测器进行压力测试的强化学习框架。StealthRL 针对检测器集成训练改述策略,同时保留语义内容,然后评估对留出检测器家族的迁移性。在完整过滤后的 MAGE 测试池(15,310 条人类样本 / 14,656 条 AI 样本)上,StealthRL 将平均...
6AI score
SaveExploits0
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
AI-text detectors face a critical robustness challenge: adversarial paraphrasing attacks that preserve semantics while evading detection. We introduce StealthRL, a reinforcement learning framework that stress-tests detector robustness under realistic adversarial conditions. StealthRL trains a...
5.5AI score
SaveExploits0
20