98 matches found
fuzzbench
FuzzBench: ファザーベンチマークのサービス FuzzBenchは、実際のベンチマークの多種多様なセットに対してファザーを評価する無料サービスであり、Google規模で動作します。FuzzBenchの目標は、ファジング研究を厳密に評価し、コミュニティがその成果を採用しやすくすることです。研究コミュニティのメンバーがファザーを提供し、評価手法の改善に関するフィードバックをいただくことを歓迎します。 FuzzBenchが提供するもの: ファザーを統合するための簡単なAPI。...
dataset
🚀 CySecBench:基于生成式AI的网络安全专用提示数据集,用于基准测试大语言模型 🛡️ 规模最大、最全面的基于生成式AI的网络安全专用数据集,用于基准测试大语言模型 🌟 概述 CySecBench 论文提供了: 🎯 一个尖端的提示数据集 ,包含12662条针对网络安全挑战的提示。 🧠 新颖的越狱方法 ,利用提示混淆与改进技术。 📊 对大语言模型(如ChatGPT、Claude和Gemini)的全面性能评估 。 为什么选择 CySecBench? 现有数据集过于宽泛,且往往缺乏对网络安全的专注。CySecBench 通过提供领域专用的提示...
wafpass
WAFPASS root@kitploit: ██╗ ██╗ █████╗ ███████╗██████╗ █████╗ ███████╗███████╗ ██║ ██║██╔══██╗██╔════╝██╔══██╗██╔══██╗██╔════╝██╔════╝ ██║ █╗ ██║███████║█████╗ ██████╔╝███████║███████╗███████╗ ██║███╗██║██╔══██║██╔══╝ ██╔═══╝ ██╔══██║╚════██║╚════██║ ╚███╔███╔╝██║ ██║██║ ██║ ██║ ██║███████║███████...
wafpass
WAFPASS root@kitploit: ██╗ ██╗ █████╗ ███████╗██████╗ █████╗ ███████╗███████╗ ██║ ██║██╔══██╗██╔════╝██╔══██╗██╔══██╗██╔════╝██╔════╝ ██║ █╗ ██║███████║█████╗ ██████╔╝███████║███████╗███████╗ ██║███╗██║██╔══██║██╔══╝ ██╔═══╝ ██╔══██║╚════██║╚════██║ ╚███╔███╔╝██║ ██║██║ ██║ ██║ ██║███████║███████...
oasis
OASIS 攻撃的AIセキュリティインテリジェンス標準 — オープンソースAIセキュリティベンチマーキング。 AIモデルが攻撃的セキュリティタスク(脆弱性発見、悪用、権限昇格など)をどの程度実行できるかをベンチマークします。完全な分析:MITRE ATT&CKマッピング、行動スコアリング、詳細レポート付き。すべてローカルで実行、自分のAPIキーを使用。アカウント不要、データはマシンから外部に出ません。 OASISを選ぶ理由...
socbench
socbench 最先端の推論LLMをSOCエージェント として、生のNetFlowデータ上でベンチマークする。...
dns-benchmark-tool
⚠️ este proyecto se ha trasladado y ya no se mantiene. instala el sucesor: pip install net-benchmark repositorio: https://github.com/net-benchmark/net-benchmark todos los comandos y opciones existentes siguen siendo completamente compatibles. DNS Benchmark Tool Parte de BuildTools - Network...
vulnerable-web-app
vulnerable-web-app ⚠️ Intentionally vulnerable. A runna...
ExploitBench AI Exploit Benchmark Tool
ExploitBench measures how far AI agents climb, from reaching vulnerable code, to triggering the bug, to building exploit primitives, to arbitrary code execution...
On the Evaluation of Spiking Neural Network Configurations for Network Intrusion Detection
Network intrusion detection is a core component of modern cybersecurity infrastructure, yet the deep learning models that dominate the field are computationally demanding, motivating interest in lightweight alternatives suited to edge and neuromorphic deployment. Spiking Neural Networks SNNs are...
How to Compare the Security of Code Written by Humans to LLM-Generated Code
Large language models LLMs are rapidly transforming how software is created and maintained. Comparing LLM-generated code against human-written standards is essential to determine whether these new tools uphold or erode the security baselines established by professional developers. Yet, we lack a...
Unlocking Apple's Private Cloud Compute: An Analysis of Privacy-Preserving Artificial Intelligence
Many existing Artificial Intelligence AI solutions on mobile devices rely on an extensive collection of sensitive data, raising privacy concerns and often requiring storage for both context and model improvement. Apple's Private Cloud Compute PCC aims to address this by emphasizing mobile device...
LLMs for Secure Hardware Design and Related Problems: Opportunities and Challenges
The integration of Large Language Models LLMs into Electronic Design Automation EDA and hardware security is rapidly reshaping the semiconductor industry. While LLMs offer unprecedented capabilities in generating Register Transfer Level RTL code, automating testbenches, and bridging the semantic...
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can gradually steer the conversation toward an unsafe answer. Recent methods demonstrate this risk, but they are usually...
Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents
Autonomous Large Language Model LLM agents are increasingly deployed to conduct complex tasks by interacting with external tools, APIs, and memory stores. However, processing untrusted external data exposes these agents to severe security threats, such as indirect prompt injection and unauthorize...
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
Large Language Model LLM agents are increasingly proposed for autonomous cybersecurity tasks, but their capabilities in realistic offensive settings remain poorly understood. We present DeepRed, an open-source benchmark for evaluating LLM-based agents on realistic Capture The Flag CTF challenges ...
HarmChip: Evaluating Hardware Security Centric LLM Safety Via Jailbreak Benchmarking
The integration of large language models LLMs into electronic design automation EDA workflows has introduced powerful capabilities for RTL generation, verification, and design optimization, but also raises critical security concerns. Malicious LLM outputs in this domain pose hardware-level threat...
n-days-poc-benchmark-and-dataset
ICS N-Day Vulnerability PoC Benchmark Suite A structured coll...
From transparency to action: What the latest Microsoft email security benchmark reveals
In our last benchmarking post, Clarity in complexity: New insights for transparent email security ,1 we shared why transparency matters more than ever in email security and how clear, consistent benchmarking helps security teams cut through noise and make confident decisions. Today, we’re...
PQC-LEO: An Evaluation Framework for Post-Quantum Cryptographic Algorithms
Advances in quantum computing threaten digital communication security by undermining the foundations of current public-key cryptography through Shor's quantum algorithm. This has driven the development of Post-Quantum Cryptography PQC, a new set of algorithms resistant to quantum attacks. While...