103 matches found
dataset
🚀 CySecBench: Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models 🛡️ The largest and most comprehensive Generative AI-based CyberSecurity-focused Dataset for Benchmarking Large Language Models 🌟 Overview The CySecBench paper offers: 🎯 A cutting-edge...
socbench
socbench Benchmark frontier reasoning LLMs as SOC agents on raw NetFlow data. socbench benchmarks frontier reasoning models as SOC agents: each model runs a bounded multi-turn agent loop against a deterministic, pre-indexed NetFlow corpus, with persona-scoped read-only tools, fixed dollar caps pe...
oasis
OASIS Offensive AI Security Intelligence Standard — Open-source AI security benchmarking. Benchmark how AI models perform offensive security tasks — vulnerability discovery, exploitation, privilege escalation, and more. Full analysis with MITRE ATT&CK mapping, behavioral scoring, and detailed...
kube-bench
kube-bench is a tool that checks whether Kubernetes is deployed securely by running the checks documented in the CIS Kubernetes Benchmark. Tests are configured with YAML files, making this tool easy to update as test specifications evolve. CIS Scanning as part of Trivy and the Trivy Operator Triv...
clusterfuzz
ClusterFuzz ClusterFuzz is a scalable fuzzing infrastructure that finds security and stability issues in software. Google uses ClusterFuzz to fuzz all Google products and as the fuzzing backend for OSS-Fuzz. ClusterFuzz provides many features which help seamlessly integrate fuzzing into a softwar...
wafpass
WAFPASS ██╗ ██╗ █████╗ ███████╗██████╗ █████╗ ███████╗███████╗ ██║ ██║██╔══██╗██╔════╝██╔══██╗██╔══██╗██╔════╝██╔════╝ ██║ █╗ ██║███████║█████╗ ██████╔╝███████║███████╗███████╗ ██║███╗██║██╔══██║██╔══╝ ██╔═══╝ ██╔══██║╚════██║╚════██║ ╚███╔███╔╝██║ ██║██║ ██║ ██║ ██║███████║███████║ ╚══╝╚══╝ ╚═╝...
wafpass
WAFPASS ██╗ ██╗ █████╗ ███████╗██████╗ █████╗ ███████╗███████╗ ██║ ██║██╔══██╗██╔════╝██╔══██╗██╔══██╗██╔════╝██╔════╝ ██║ █╗ ██║███████║█████╗ ██████╔╝███████║███████╗███████╗ ██║███╗██║██╔══██║██╔══╝ ██╔═══╝ ██╔══██║╚════██║╚════██║ ╚███╔███╔╝██║ ██║██║ ██║ ██║ ██║███████║███████║ ╚══╝╚══╝ ╚═╝...
oss-fuzz-gen
퍼즈 타깃 생성 및 평가를 위한 프레임워크 이 프레임워크는 다양한 대규모 언어 모델LLM을 사용하여 실제 C/C++/Java/Python 프로젝트를 위한 퍼즈 타깃을 생성하고, OSS-Fuzz 플랫폼을 통해 벤치마킹합니다. 자세한 내용은 AI 기반 퍼징: 버그 사냥 장벽을 허무는 방법에서 확인할 수 있습니다: 현재 지원되는 모델은 다음과 같습니다: Vertex AI code-bison Vertex AI code-bison-32k Gemini Pro Gemini Ultra Gemini Experimental Gemini 1.5...
fuzzbench
FuzzBench: 퍼저 벤치마킹 서비스 FuzzBench는 실제 세계의 다양한 벤치마크에서 퍼저를 평가하는 무료 서비스로, Google 규모로 운영됩니다. FuzzBench의 목표는 퍼징 연구를 엄격하게 평가하는 과정을 간편하게 만들고, 커뮤니티가 퍼징 연구를 더 쉽게 채택할 수 있도록 하는 것입니다. 연구 커뮤니티 구성원들이 자신의 퍼저를 기여하고 평가 기술 개선에 대한 피드백을 제공해 주시기 바랍니다. FuzzBench는 다음을 제공합니다: 퍼저 통합을 위한 간편한 API 실제 프로젝트의 벤치마크. FuzzBench는 모...
cleverhans
CleverHans 최신 릴리스: v4.0.0 이 저장소는 적대적 예제adversarial examples에 대한 머신러닝 시스템의 취약성을 벤치마킹하기 위한 Python 라이브러리인 CleverHans의 소스 코드를 포함합니다. 이러한 취약성에 대해 더 자세히 알아보려면 관련 블로그를 참조하세요. CleverHans 라이브러리는 지속적으로 개발 중이며, 최신 공격 및 방어 기법의 기여contributions를 항상 환영합니다. 특히, 현재 열려 있는 이슈issues를 해결하는 데 도움을 주시면 항상 환영합니다. v4.0.0부...
dns-benchmark-tool
⚠️ this project has moved and is no longer maintained. install the successor: pip install net-benchmark repository: https://github.com/net-benchmark/net-benchmark all existing commands and flags remain fully compatible. DNS Benchmark Tool Part of BuildTools - Network Performance Suite Fast,...
MADBench: Benchmarking the Security of Multi-Agent Debate
Multi-agent debate MAD can improve large language model LLM reasoning by allowing multiple agents to exchange and critique their answers to the same task. However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect...
Awesome-LLMs-for-Vulnerability-Detection — Updated!
Awesome Large Language Models for Vulnerability Detection A curated list of papers, projects, and agent skills on using LLMs for vulnerability detection and discovery. 📄 Papers Only showing 2025 and later. For earlier work, see Papers Archive 2024 and earlier. Title| Venue| Year| Paper| Github...
vulnerable-web-app
vulnerable-web-app ⚠️ Intentionally vulnerable. A runna...
ExploitBench AI Exploit Benchmark Tool
ExploitBench measures how far AI agents climb, from reaching vulnerable code, to triggering the bug, to building exploit primitives, to arbitrary code execution...
On the Evaluation of Spiking Neural Network Configurations for Network Intrusion Detection
Network intrusion detection is a core component of modern cybersecurity infrastructure, yet the deep learning models that dominate the field are computationally demanding, motivating interest in lightweight alternatives suited to edge and neuromorphic deployment. Spiking Neural Networks SNNs are...
How to Compare the Security of Code Written by Humans to LLM-Generated Code
Large language models LLMs are rapidly transforming how software is created and maintained. Comparing LLM-generated code against human-written standards is essential to determine whether these new tools uphold or erode the security baselines established by professional developers. Yet, we lack a...
Unlocking Apple's Private Cloud Compute: An Analysis of Privacy-Preserving Artificial Intelligence
Many existing Artificial Intelligence AI solutions on mobile devices rely on an extensive collection of sensitive data, raising privacy concerns and often requiring storage for both context and model improvement. Apple's Private Cloud Compute PCC aims to address this by emphasizing mobile device...
LLMs for Secure Hardware Design and Related Problems: Opportunities and Challenges
The integration of Large Language Models LLMs into Electronic Design Automation EDA and hardware security is rapidly reshaping the semiconductor industry. While LLMs offer unprecedented capabilities in generating Register Transfer Level RTL code, automating testbenches, and bridging the semantic...
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can gradually steer the conversation toward an unsafe answer. Recent methods demonstrate this risk, but they are usually...