836 matches found
morty
Morty Web content sanitizer proxy as a service Morty rewrites web pages to exclude malicious HTML tags and attributes. It also replaces external resource references to prevent third party information leaks. The main goal of morty is to provide a result proxy for searx, but it can be used as a...
cve-bench
CVE-Bench A benchmark for evaluating LLM agents on fixing real-world security vulnerabilities. Agents run inside sandboxed Docker containers and are scored against the maintainer's security test suite. Requirements Python 3.12+ Docker OPENAIAPIKEY, ANTHROPICAPIKEY, and/or POOLSIDEAPIKEY in your...
MalEval
MalEval Article: Is “Knowing It’s Malicious” Enough? Evaluating LLMs for Fine-Grained Malware Behavior Auditing Article DOI: 10.1145/3832187 MalEval is a framework for evaluating Android malware behavior reports generated by large language models. The code in this repository implements two...
terraform-aws-secure-baseline
terraform-aws-secure-baseline Terraform Module Registry A terraform module to set up your AWS account with the reasonably secure configuration baseline. Most configurations are based on CIS Amazon Web Services Foundations v1.4.0 and AWS Foundational Security Best Practices v1.0.0. See Benchmark...
terraform-aws-secure-baseline
terraform-aws-secure-baseline Terraform Module Registry Un módulo de Terraform para configurar tu cuenta de AWS con una línea base de seguridad razonablemente segura. La mayoría de las configuraciones se basan en CIS Amazon Web Services Foundations v1.4.0 y AWS Foundational Security Best Practice...
delirium-ai-safety-benchmark
Delirium AI Safety Benchmark A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion ACE and related liminal attack vectors. Delirium is not an exploitation tool. It is a standardized benchmark designed to detect the precise moment when a language model's attention...
prompt_injection
Cross-Surface Prompt Injection Benchmark Benchmark for measuring where in a tool-using LLM agent's execution pipeline a prompt injection defense fires — and where it doesn't. Embeds unique canary tokens SECRET-A-F0-98 in injected payloads and tracks them at four pipeline stages: exposed → persist...
vulnrepro-benchmark
VulnRepro A benchmark that checks if an AI model can actually review vulnerable code, or if it just sounds confident. Most security benchmarks ask one question: can the model find the bug? That is only half the job. The other half, the part that actually wears you down in real review work, is not...
Zeus
Zeus AWS Auditing & Hardening Tool Zeus is a powerful tool for AWS EC2 / S3 / CloudTrail / CloudWatch / KMS best hardening practices. It checks security settings according to the profiles the user creates and changes them to recommended settings based on the CIS AWS Benchmark source at request of...
rp
rp++: a fast ROP gadget finder for PE/ELF/Mach-O x86/x64/ARM/ARM64 binaries Overview rp++ or rp is a C++ ROP gadget finder for PE/ELF/Mach-O executables and x86/x64/ARM/ARM64 architectures. Finding ROP gadgets To find ROP gadget you need to specify a file with the --file / -f option and use the...
whalescan
Whalescan Escáner de vulnerabilidades para contenedores de Windows. Primeros pasos git clone https://github.com/saira-h/whalescan pip install -r requirements.txt python main.py Descripción general Whalescan realiza varias comprobaciones de referencia, así como la verificación de CVEs. Esta...
FinRED-paper
FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming IEEE ICDM 2026 A red-team benchmark generation pipeline for safety evaluation in the financial domain. Supplementary Documentation Detailed materials referenced in the paper:...
CTFTiny
CTFTiny: Lite Benchmarking Offensive Cyber Skills in Large Language Models This is the official repository for CTFTiny from "Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark" AAAI'26 paper. For CTFJudge, please refer to CTFJud...
docker-bench-security
Docker Bench for Security The Docker Bench for Security is a script that checks for dozens of common best-practices around deploying Docker containers in production. The tests are all automated, and are based on the CIS Docker Benchmark v1.6.0. We are making this available as an open-source utili...
robustbench
RobustBench: a standardized adversarial robustness benchmark Francesco Croce University of Tübingen, Maksym Andriushchenko EPFL, Vikash Sehwag Princeton University, Edoardo Debenedetti EPFL, Nicolas Flammarion EPFL, Mung Chiang Purdue University, Prateek Mittal Princeton University, Matthias Hein...
chain-bench
📖 Documentacióndocs Chain-bench es una herramienta de código abierto para auditar su pila de cadena de suministro de software en busca de cumplimiento de seguridad basada en un nuevo punto de referencia CIS para la cadena de suministro de software. La auditoría se centra en todo el proceso del...
terminal-bench-2
Terminal-Bench 2.0 | | | | || || | |/ \ '| ' | | ' \ / | | || || | | / | | | | | | | | | | | | | | || || |||| || || |||| ||,|| |||| || | | | | \ / \ \\ | \ / \ ' \ / | ' \ || | | | \\ | | | / | | | | | | | / / | || | \ \ |/ || |||| || |/ \\ Terminal-Bench is a popular benchmark for...
legba
legba Join the project community on our server! Legba is a multiprotocol credentials bruteforcer / password sprayer and enumerator built with Rust and the Tokio asynchronous runtime in order to achieve better performances and stability while consuming less resources than similar tools. Key Featur...
FML-Network
FLNET2023: Realistic Network Intrusion Detection Dataset for Federated Learning Paper: FLNET2023: Realistic Network Intrusion Detection Dataset for Federated Learning Dataset: FLNET2023 Introduction FLNET2023 is a state-of-the-art benchmark dataset for intrusion detection systems, specifically fo...
trivy-operator
Kubernetes-native security toolkit. Documentation Introduction The Trivy Operator leverages Trivy to continuously scan your Kubernetes cluster for security issues. The scans are summarised in security reports as Kubernetes Custom Resource Definitions, which become accessible through the Kubernete...