1 matches found
Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing
Recent autonomous penetration testing papers report high benchmark scores while adding multi-component security harnesses around frontier LLMs. Because these systems often change both architecture and backbone model, it is difficult to tell how much performance comes from the harness rather than...
6.1AI score
SaveExploits0
20