1 matches found
When Context Gets Root: Privilege Escalation in LLM Harnesses
Instruction hierarchy is a model-side defense that assigns instructions different levels of privilege according to their sources. These levels constrain which content may direct model behavior. During agent execution, however, agent harnesses construct context for each model invocation. This...
6.2AI score
SaveExploits0
20