Lucene search
+L

5 matches found

Packet Storm News
Packet Storm News
•added 2026/09/08 12:00 a.m.•32 views

Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks

Large language model LLM benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cybersecurity benchmarks across 10 proprietary, open-weight, and cybersecurity-specialized LLMs. By modeling benchmarks as...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/06/22 12:00 a.m.•15 views

RIFT-Bench: Dynamic Red-Teaming for Agentic AI Systems

Agentic AI systems powered by large language models LLMs are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities. Existing security evaluations are often tied to specific implementations or domains, limiting unified...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/06/03 12:00 a.m.•263 views

CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-To-End Cybersecurity Capabilities

AI has the potential to transform cybersecurity by enabling systems that can autonomously detect, analyze, and remediate software vulnerabilities. However, existing cybersecurity evaluations of AI systems are limited in scale or scope, and fail to capture the end-to-end lifecycle of real-world...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/02/22 12:00 a.m.•14 views

Red-Teaming Claude Opus and ChatGPT-Based Security Advisors for Trusted Execution Environments

Trusted Execution Environments TEEs e.g., Intel SGX and ArmTrustZone aim to protect sensitive computation from a compromised operating system, yet real deployments remain vulnerable to microarchitectural leakage, side-channel attacks, and fault injection. In parallel, security teams increasingly...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/05/28 12:00 a.m.•12 views

GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance

DNA, encoding genetic instructions for almost all living organisms, fuels groundbreaking advances in genomics and synthetic biology. Recently, DNA Foundation Models have achieved success in designing synthetic functional DNA sequences, even whole genomes, but their susceptibility to jailbreaking...

7.5AI score
SaveExploits0
Rows per page
Query Builder