1 matches found
HOL Guard AI Agent Runtime Security Benchmark
HOL Guard Benchmark is an open source fixture-based benchmark for evaluating security controls applied to AI coding agents. It defines deterministic scenarios covering secret access, risky shell execution, tool poisoning, approval handling, safe operations, and receipt generation across Codex CLI...
5.8AI score
SaveExploits0
20