2 matches found
benign-instruction-bench
Passing the Test You Trained On Re-evaluating prompt-injection detectors where LLM agents actually use them: on the tool outputs an agent reads. Teams pick injection detectors by their benchmark scores. We check whether those scores predict behavior inside an agent, and find that they mostly...
SaveExploits0
Effective Red-Teaming of Policy-Adherent Agents
Task-oriented LLM-based agents are increasingly used in domains with strict policies, such as refund eligibility or cancellation rules. The challenge lies in ensuring that the agent consistently adheres to these rules and policies, appropriately refusing any request that would violate them, while...
6.9AI score
SaveExploits0
20