1 matches found
Evaluating Coding Agents on Kernel Exploit Generation
Coding agents now find real vulnerabilities in production software. However, bug discovery results do not measure whether agents can construct exploit primitives. We introduce KEX-bench, a benchmark for evaluating coding agents on exploit primitive generation against real operating-system kernels...
6AI score
SaveExploits0
20