2 matches found
terminal-bench-2
Terminal-Bench 2.0 root@kitploit: | | | | || || | |/ \ '| ' | | ' \ / | | || || | | / | | | | | | | | | | | | | | || || |||| || || |||| ||,|| |||| || | | | | \ / \ \\ | \ / \ ' \ / | ' \ || | | | \\ | | | / | | | | | | | / / | || | \ \ |/ || |||| || |/ \\...
5.3AI score
SaveExploits0
Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops
Agent benchmarks score submissions with outcome verifiers that are typically hand-written and brittle, leaving them open to reward hacking. We audit 1,968 tasks across five terminal-agent benchmarks and find 323 16% hackable by frontier models given only the task description. This corrupts both...
5.5AI score
SaveExploits0
20