1 matches found
CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents
Personalized language agents use persistent memory to adapt to users over time, but the same mechanism creates an attack surface. When new information conflicts with stored preferences, an agent must distinguish genuine preference drift from temporary context shifts, ambiguity, or adversarial...
5.9AI score
SaveExploits0
20