Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2026/04/21 12:0 a.m.12 views

Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges

Large Language Model LLM agents are increasingly proposed for autonomous cybersecurity tasks, but their capabilities in realistic offensive settings remain poorly understood. We present DeepRed, an open-source benchmark for evaluating LLM-based agents on realistic Capture The Flag CTF challenges ...

6AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/07/05 12:0 a.m.7 views

Can Large Language Models Automate the Refinement of Cellular Network Specifications?

Cellular networks serve billions of users globally, yet concerns about reliability and security persist due to weaknesses in 3GPP standards. However, traditional analysis methods, including manual inspection and automated tools, struggle with increasingly expanding cellular network specifications...

6.9AI score
SaveExploits0
Rows per page
Query Builder