Lucene search
+L

737 matches found

Kitploit
Kitploit
•added 2026/10/03 9:05 a.m.•16 views

Agentic-CLIP-Benchmark

Agentic-CLIP-Benchmark: Evaluación Zero-Shot en CIFAR-10 -green.svg Este proyecto implementa un pipeline robusto y automatizado para evaluar el modelo CLIP ViT-B/32 de OpenAI en el conjunto de prueba completo de CIFAR-10 10,000 imágenes. Desarrollado mediante un flujo de trabajo AI-Native Trae ID...

9.8CVSS6.3AI score0.02187EPSS
SaveExploits1References1
Packet Storm News
Packet Storm News
•added 2026/10/03 12:00 a.m.•2 views

COPEX: Benchmarking LLM Robustness to Adversarial Context across Model Context Protocol Layers

Large language models increasingly mediate tool use in Model Context Protocol MCP systems, where adversarial influence may enter through user instructions, tool schemas, tool outputs, or protocol messages. Existing benchmarks often evaluate deployed agents, conflating model susceptibility with...

SaveExploits0
Chainguard
Chainguard
•added 2026/10/02 8:54 p.m.•14 views

GO-2026-6505 vulnerabilities

Vulnerabilities for packages: cloud-provider-gcp-cloud-controller-manager-fips, grype-fips, crossplane-cli-fips, scorecard, terraform-provider-google, aws-fsx-csi-driver-fips, cilium, k6-operator, knative-serving, cert-manager-istio-csr-fips, kubernetes-csi-driver-nfs-fips, gitlab-kas,...

5.8AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/10/02 7:43 p.m.•18 views

CUAHarm

Medición de la nocividad de los agentes que usan computadoras 🤗 Hugging Face 📄 Paper 📋 Introducción CUAHarm es un punto de referencia diseñado para evaluar los riesgos de seguridad de los Agentes que Usan Computadoras CUAs, por sus siglas en inglés: agentes de IA que pueden controlar computadoras...

6.3AI score
SaveExploits0References5
Kitploit
Kitploit
•added 2026/10/02 1:48 p.m.•4 views

MEA-Bench

MEA-Bench: Un Benchmark para Ataques de Extracción de Modelos Este repositorio proporciona un benchmark unificado para ataques de extracción de modelos, defensas, ataques adaptativos y evaluación. La interfaz pública está organizada en torno a un pequeño número de comandos portables. Los módulos...

6.2AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/02 9:07 a.m.•17 views

windows_hardening

HardeningKitty y el endurecimiento de Windows Introducción El proyecto comenzó como una simple lista de endurecimiento para Windows 10. Después de un tiempo, se creó HardeningKitty para simplificar el endurecimiento de Windows. Ahora, HardeningKitty admite las guías de Microsoft, CIS Benchmarks,...

9CVSS8.1AI score0.99792EPSS
SaveExploits42References2
Kitploit
Kitploit
•added 2026/10/02 7:36 a.m.•12 views

TarantuBench

TarantuBench v1 Un benchmark para evaluar agentes de IA en desafíos de seguridad web, generado por el motor de TarantuLabs. ¿Qué es esto? TarantuBench es una colección de 100 aplicaciones web vulnerables, cada una con una bandera oculta TARANTU.... La tarea de un agente es encontrar y extraer la...

6.2AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/10/02 12:51 a.m.•210 views

gdbfuzz

GDBFuzz: Debugger-Driven Fuzzing This is the companion code for the paper: 'Fuzzing Embedded Systems using Debugger Interfaces'. A preprint of the paper can be found here https://publications.cispa.saarland/3950/. The code allows the users to reproduce and extend the results reported in the paper...

6AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/10/02 12:25 a.m.•14 views

dns-benchmark-tool

⚠️ 此项目已迁移,不再维护。 请安装替代项目:pip install net-benchmark 仓库:https://github.com/net-benchmark/net-benchmark 所有现有命令和标志完全兼容。 DNS 基准测试工具 属于 BuildTools - 网络性能套件 快速、全面的 DNS 性能测试,支持 DNSSEC 验证、DoH/DoT 及企业级功能 bash pip install dns-benchmark-tool dns-benchmark benchmark --use-defaults --formats csv,excel 🎉 1,400+...

6.3AI score
SaveExploits0References9
Kitploit
Kitploit
•added 2026/10/01 7:17 p.m.•12 views

FuzzingBrain-Bench

FuzzingBrain Bench Un benchmark para la reproducción de vulnerabilidades impulsada por LLM en 77 errores de día cero reales en 43 proyectos de código abierto C / C++ / Java. Cada desafío le da al agente solo el harness de fuzzing el objetivo y el código fuente del proyecto en la revisión vulnerab...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/01 3:37 p.m.•10 views

WAInjectBench

WAInjectBench WAInjectBench es un benchmark integral para la detección de inyección de prompts en agentes web. Cubre 6 tipos de ataques , en dos modalidades : texto e imagen. 📂 Estructura del conjunto de datos data/ text/ benign/ → 4 categorías, almacenadas como archivos JSONL malicious/ → 8 tipo...

6.3AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/10/01 2:46 p.m.•5 views

prompt_injection

Benchmark de inyección de prompts entre superficies Benchmark para medir en qué punto del pipeline de ejecución de un agente LLM que usa herramientas se activa una defensa contra inyección de prompts — y en qué punto no. Incrusta tokens canary únicos SECRET-A-F0-98 en payloads inyectados y los...

6.3AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/10/01 12:11 p.m.•12 views

delirium-ai-safety-benchmark

Delirium: Benchmark de Seguridad de IA Un marco de diagnóstico para medir la vulnerabilidad de los LLM ante la Erosión Contextual Afectiva ACE y vectores de ataque liminales relacionados. Delirium no es una herramienta de explotación. Es un benchmark estandarizado diseñado para detectar el moment...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/01 9:42 a.m.•7 views

FinRED-paper

FinRED: Conjunto de datos de evaluación de red-teaming financiero Una canalización de generación de benchmarks de red-team para la evaluación de seguridad en el dominio financiero. Documentación complementaria Materiales detallados referenciados en el artículo: docs/expertvalidation.md — Resultad...

6.3AI score
SaveExploits0References4
Packet Storm News
Packet Storm News
•added 2026/10/01 12:00 a.m.•11 views

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable...

6AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/09/30 10:35 p.m.•12 views

Ajar

Ajar: Measuring Open Privilege in Agent Defenses An agent-security benchmark reports two numbers, attack success and benign utility, and both are read off runs that happened. Neither says what the defense stood ready to allow on the paths no run took. Ajar asks it directly: for each benign task i...

6.3AI score
SaveExploits0References1
Packet Storm News
Packet Storm News
•added 2026/09/30 12:00 a.m.•10 views

Do Defenses against LLM Extraction Work across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction

Large language models LLMs deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attacks and defenses, evaluations remain fragmented across access...

5.9AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/09/27 3:23 p.m.•21 views

ActBench

ActBench ActBench is a self-evolving benchmark of behavioral safety in cowork agents. It defines behavioral safety as whether an agent's execution remains within the permissions and state changes required by a benign task, and evaluates realized behavioral risk from execution trajectories rather...

6.2AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/09/25 5:24 p.m.•12 views

ethibench

From Controlled to the Wild: Evaluation of Pentesting Agents in the Real-World AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which systems will perform best on real-world targets. Most existing evaluations...

6.4AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/09/14 10:16 p.m.•20 views

vader

VADER:用于漏洞评估、检测、解释和修复的人工评估基准 官方 GitHub 仓库:https://github.com/AfterQuery/vader Hugging Face 数据集:https://huggingface.co/datasets/AfterQuery/vader VADER 是一个人工评估基准 ,旨在衡量大型语言模型(LLMs)处理真实世界软件漏洞的能力。它包含 174 个真实世界漏洞案例 (从开源仓库中精选),涵盖四项任务: 漏洞识别与分类(CWE) 根因解释 补丁(修复) 测试计划生成 这些案例涵盖 15+ 编程语言 (例如...

6AI score
SaveExploits0References1
Rows per page
Query Builder