Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2025/06/22 12:0 a.m.34 views

SEC-Bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks

Rigorous security-focused evaluation of large language model LLM agents is imperative for establishing trust in their safe deployment throughout the software development lifecycle. However, existing benchmarks largely rely on synthetic challenges or simplified vulnerability datasets that fail to...

7.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/06 12:0 a.m.6 views

Benchmarking Misuse Mitigation against Covert Adversaries

Existing language model safety evaluations focus on overt attacks and low-stakes tasks. Realistic attackers can subvert current safeguards by requesting help on small, benign-seeming tasks across many independent queries. Because individual queries do not appear harmful, the attack is hard to...

7.2AI score
SaveExploits0
Rows per page
Query Builder