1 matches found
AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents
Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion AgentBench or security robustness AgentDojo, ASB, rather than the complete pipeline of planning, tool selection, tool execution, memory and reasoning. Failures can occur at any stage, yet existing...
6AI score
SaveExploits0
20