5 matches found
Rubrics-as-an-Attack-Surface
Rúbricas como superficie de ataque: Deriva de preferencias sigilosa en jueces LLM 📊 Dataset • 🤖 Modelos entrenados • 📝 Artículo • 💻 Repositorio Este repositorio contiene el código del artículo Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges de Ruomeng Ding, Yifei Pang, He Su...
Attack Success Rate Is Not a Number: On Measurement Validity in Agentic AI Security Evaluation
Attack success rate ASR is the headline metric in nearly every published evaluation of attacks on, and defenses for, LLM agents. We argue that ASR as currently used is not a single quantity but a family of metrics parameterized by six design choices that papers seldom specify and never hold...
Ranked by the Matcher: A Reproducibility Audit of Knowledge Graph Extraction from Threat Reports
Security teams and researchers choose knowledge-graph extraction tooling for threat reports on the strength of published triple-F1 scores, yet those scores depend on how predicted triples are matched to gold annotations. We could reimplement the stated matching rule for only five of twelve...
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous...
Adversarial Attacks on LLM-As-A-Judge Systems: Insights from Prompt Injections
LLM as judge systems used to assess text quality code correctness and argument strength are vulnerable to prompt injection attacks. We introduce a framework that separates content author attacks from system prompt attacks and evaluate five models Gemma 3.27B Gemma 3.4B Llama 3.2 3B GPT 4 and Clau...