1 matches found
Compromise Is Not Consequence: Evaluating Task-Scoped Authorization in LLM Agents with Paired Replay
A tool-using model can follow a malicious instruction even when its credentials are valid. We study whether task-scoped authorization contains the resulting tool execution. Our paired-replay testbed samples a model request once and submits the same action, resource, and arguments to broad bearer,...
6AI score
SaveExploits0
20