2 matches found
Rubrics-as-an-Attack-Surface
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges π Dataset β’ π€ Trained Models β’ π Paper β’ π» Repo This repository contains code for the paper Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges by Ruomeng Ding, Yifei Pang, He Sun, Yizhong Wang, Steven Wu, and...
6.5AI score
SaveExploits0References1
CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents
Personalized language agents use persistent memory to adapt to users over time, but the same mechanism creates an attack surface. When new information conflicts with stored preferences, an agent must distinguish genuine preference drift from temporary context shifts, ambiguity, or adversarial...
5.9AI score
SaveExploits0
20