Lucene search
+L

2 matches found

Kitploit
Kitploit
β€’added 2026/10/06 4:55 a.m.β€’17 views

Rubrics-as-an-Attack-Surface

Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges πŸ“Š Dataset β€’ πŸ€– Trained Models β€’ πŸ“ Paper β€’ πŸ’» Repo This repository contains code for the paper Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges by Ruomeng Ding, Yifei Pang, He Sun, Yizhong Wang, Steven Wu, and...

6.5AI score
SaveExploits0References1
Packet Storm News
Packet Storm News
β€’added 2026/09/02 12:00 a.m.β€’51 views

CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents

Personalized language agents use persistent memory to adapt to users over time, but the same mechanism creates an attack surface. When new information conflicts with stored preferences, an agent must distinguish genuine preference drift from temporary context shifts, ambiguity, or adversarial...

5.9AI score
SaveExploits0
Rows per page
Query Builder