2 matches found
vader
VADER:用于漏洞评估、检测、解释和修复的人工评估基准 官方 GitHub 仓库:https://github.com/AfterQuery/vader Hugging Face 数据集:https://huggingface.co/datasets/AfterQuery/vader VADER 是一个人工评估基准 ,旨在衡量大型语言模型(LLMs)处理真实世界软件漏洞的能力。它包含 174 个真实世界漏洞案例 (从开源仓库中精选),涵盖四项任务: 漏洞识别与分类(CWE) 根因解释 补丁(修复) 测试计划生成 这些案例涵盖 15+ 编程语言 (例如...
6AI score
SaveExploits0References1
LLM-GUARD: Large Language Model-Based Detection and Repair of Bugs and Security Vulnerabilities in C++ and Python
Large Language Models LLMs such as ChatGPT-4, Claude 3, and LLaMA 4 are increasingly embedded in software/application development, supporting tasks from code generation to debugging. Yet, their real-world effectiveness in detecting diverse software bugs, particularly complex, security-relevant...
7.2AI score
SaveExploits0
20