Lucene search
+L

4 matches found

Kitploit
Kitploit
•added 2026/10/01 7:17 p.m.•12 views

FuzzingBrain-Bench

FuzzingBrain Bench A benchmark for LLM-driven vulnerability reproduction on 77 real zero-day bugs across 43 open-source projects C / C++ / Java. Each challenge gives the agent only the fuzz harness the target and the project source at the vulnerable revision — no patch, no fix commit, no target...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/01 12:11 p.m.•12 views

delirium-ai-safety-benchmark

Delirium AI Safety Benchmark A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion ACE and related liminal attack vectors. Delirium is not an exploitation tool. It is a standardized benchmark designed to detect the precise moment when a language model's attention...

6.3AI score
SaveExploits0References1
Packet Storm News
Packet Storm News
•added 2026/10/01 12:00 a.m.•10 views

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable...

6AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/09/14 10:16 p.m.•20 views

vader

VADER:用于漏洞评估、检测、解释和修复的人工评估基准 官方 GitHub 仓库:https://github.com/AfterQuery/vader Hugging Face 数据集:https://huggingface.co/datasets/AfterQuery/vader VADER 是一个人工评估基准 ,旨在衡量大型语言模型(LLMs)处理真实世界软件漏洞的能力。它包含 174 个真实世界漏洞案例 (从开源仓库中精选),涵盖四项任务: 漏洞识别与分类(CWE) 根因解释 补丁(修复) 测试计划生成 这些案例涵盖 15+ 编程语言 (例如...

6AI score
SaveExploits0References1
Rows per page
Query Builder