103 matches found
clusterfuzz
ClusterFuzz ClusterFuzz is a scalable fuzzing infrastructure that finds security and stability issues in software. Google uses ClusterFuzz to fuzz all Google products and as the fuzzing backend for OSS-Fuzz. ClusterFuzz provides many features which help seamlessly integrate fuzzing into a softwar...
fuzzbench
FuzzBench: Evaluación Comparativa de Fuzzers como Servicio FuzzBench es un servicio gratuito que evalúa fuzzers en una amplia variedad de referencias del mundo real, a escala de Google. El objetivo de FuzzBench es facilitar la evaluación rigurosa de la investigación en fuzzing y hacer que dicha...
dataset
🚀 CySecBench: Dataset de Prompts de Ciberseguridad Centrado en IA Generativa para la Evaluación Comparativa de Grandes Modelos de Lenguaje 🛡️ El dataset más grande y completo de Ciberseguridad basado en IA Generativa para la Evaluación Comparativa de Grandes Modelos de Lenguaje 🌟 Descripción...
socbench
socbench Evalúa modelos de razonamiento de frontera como agentes SOC sobre datos NetFlow en bruto. socbench evalúa modelos de razonamiento de frontera como agentes SOC: cada modelo ejecuta un bucle de agente multi-turno acotado contra un corpus NetFlow determinista e indexado previamente, con...
oasis
OASIS Estándar de Inteligencia de Seguridad Ofensiva con IA — Evaluación comparativa de seguridad con IA de código abierto. Evalúa cómo los modelos de IA realizan tareas de seguridad ofensiva: descubrimiento de vulnerabilidades, explotación, escalada de privilegios y más. Análisis completo con...
kube-bench
kube-bench es una herramienta que verifica que Kubernetes esté desplegado de forma segura ejecutando las comprobaciones documentadas en el CIS Kubernetes Benchmark. Las pruebas se configuran con archivos YAML, lo que facilita la actualización de la herramienta a medida que evolucionan las...
cleverhans
CleverHans latest release: v4.0.0 This repository contains the source code for CleverHans, a Python library to benchmark machine learning systems' vulnerability to adversarial examples. You can learn more about such vulnerabilities on the accompanying blog. The CleverHans library is under continu...
oss-fuzz-gen
A Framework for Fuzz Target Generation and Evaluation This framework generates fuzz targets for real-world C/C++/Java/Python projects with various Large Language Models LLM and benchmarks them via the OSS-Fuzz platform. More details available in AI-Powered Fuzzing: Breaking the Bug Hunting Barrie...
wafpass
WAFPASS ██╗ ██╗ █████╗ ███████╗██████╗ █████╗ ███████╗███████╗ ██║ ██║██╔══██╗██╔════╝██╔══██╗██╔══██╗██╔════╝██╔════╝ ██║ █╗ ██║███████║█████╗ ██████╔╝███████║███████╗███████╗ ██║███╗██║██╔══██║██╔══╝ ██╔═══╝ ██╔══██║╚════██║╚════██║ ╚███╔███╔╝██║ ██║██║ ██║ ██║ ██║███████║███████║ ╚══╝╚══╝ ╚═╝...
wafpass
WAFPASS ██╗ ██╗ █████╗ ███████╗██████╗ █████╗ ███████╗███████╗ ██║ ██║██╔══██╗██╔════╝██╔══██╗██╔══██╗██╔════╝██╔════╝ ██║ █╗ ██║███████║█████╗ ██████╔╝███████║███████╗███████╗ ██║███╗██║██╔══██║██╔══╝ ██╔═══╝ ██╔══██║╚════██║╚════██║ ╚███╔███╔╝██║ ██║██║ ██║ ██║ ██║███████║███████║ ╚══╝╚══╝ ╚═╝...
dns-benchmark-tool
⚠️ este proyecto se ha trasladado y ya no se mantiene. instala el sucesor: pip install net-benchmark repositorio: https://github.com/net-benchmark/net-benchmark todos los comandos y opciones existentes siguen siendo completamente compatibles. DNS Benchmark Tool Parte de BuildTools - Network...
MADBench: Benchmarking the Security of Multi-Agent Debate
Multi-agent debate MAD can improve large language model LLM reasoning by allowing multiple agents to exchange and critique their answers to the same task. However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect...
Awesome-LLMs-for-Vulnerability-Detection — Updated!
Awesome Large Language Models for Vulnerability Detection A curated list of papers, projects, and agent skills on using LLMs for vulnerability detection and discovery. 📄 Papers Only showing 2025 and later. For earlier work, see Papers Archive 2024 and earlier. Title| Venue| Year| Paper| Github...
vulnerable-web-app
vulnerable-web-app ⚠️ Intentionally vulnerable. A runna...
ExploitBench AI Exploit Benchmark Tool
ExploitBench measures how far AI agents climb, from reaching vulnerable code, to triggering the bug, to building exploit primitives, to arbitrary code execution...
On the Evaluation of Spiking Neural Network Configurations for Network Intrusion Detection
Network intrusion detection is a core component of modern cybersecurity infrastructure, yet the deep learning models that dominate the field are computationally demanding, motivating interest in lightweight alternatives suited to edge and neuromorphic deployment. Spiking Neural Networks SNNs are...
How to Compare the Security of Code Written by Humans to LLM-Generated Code
Large language models LLMs are rapidly transforming how software is created and maintained. Comparing LLM-generated code against human-written standards is essential to determine whether these new tools uphold or erode the security baselines established by professional developers. Yet, we lack a...
Unlocking Apple's Private Cloud Compute: An Analysis of Privacy-Preserving Artificial Intelligence
Many existing Artificial Intelligence AI solutions on mobile devices rely on an extensive collection of sensitive data, raising privacy concerns and often requiring storage for both context and model improvement. Apple's Private Cloud Compute PCC aims to address this by emphasizing mobile device...
LLMs for Secure Hardware Design and Related Problems: Opportunities and Challenges
The integration of Large Language Models LLMs into Electronic Design Automation EDA and hardware security is rapidly reshaping the semiconductor industry. While LLMs offer unprecedented capabilities in generating Register Transfer Level RTL code, automating testbenches, and bridging the semantic...
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can gradually steer the conversation toward an unsafe answer. Recent methods demonstrate this risk, but they are usually...