1 matches found
The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions
Agent guardrails are checks that approve or refuse each action before an LLM executes it. Sometimes they refuse requests that are genuinely safe. This over-safety blocks deployment when a guardrail refuses an authorized task. Evaluating over-safety is hard: at the boundary an authorized action...
5.9AI score
SaveExploits0
20