1 matches found
Efficient Auditing of Adversarial AI Agent Behavior from Agent Traces
AI agents powered by large language models LLMs can perform complex tasks but may harm the systems they operate in, either intentionally or unintentionally. Existing agent monitoring approaches rely on rule-based guardrails or LLM-based trace auditing. However, rule-based guardrails can be bypass...
5.9AI score
SaveExploits0
20