Lucene search
+L

11 matches found

Kitploit
Kitploit
•added 2026/10/06 3:57 p.m.•18 views

reverse-captcha-eval

Reverse CAPTCHA: Evaluating LLM Susceptibility to Invisible Unicode Instruction Injection An evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in otherwise normal-looking text. Where traditional CAPTCHAs exploit tasks humans can...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/06 8:35 a.m.•20 views

MalEval

MalEval Artículo: ¿Es “saber que es malicioso” suficiente? Evaluación de LLMs para la auditoría detallada del comportamiento de malware DOI del artículo: 10.1145/3832187 MalEval es un marco para evaluar informes de comportamiento de malware Android generados por grandes modelos de lenguaje. El...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/06 8:06 a.m.•13 views

CTFTiny

CTFTiny: Evaluación comparativa ligera de habilidades ofensivas en ciberseguridad de grandes modelos de lenguaje Este es el repositorio oficial de CTFTiny del artículo "Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark" AAAI'26...

6.1AI score
SaveExploits0References2
Kitploit
Kitploit
•added 2026/10/06 7:59 a.m.•20 views

claude_opus_cve_2023_0266

Demonstration that Claude 3 Opus does not understand CVE-2023-0266 and does not find it Demo 1. Even if told where the bug is Opus does not find it, and hallucinates the presence of lock acquisitions Demo 2. "Prompt engineering" aka telling the LLM exactly how to find the bug also doesn't work De...

7.9CVSS7AI score0.03702EPSS
SaveExploits0
Kitploit
Kitploit
•added 2026/10/06 6:59 a.m.•19 views

bloom

Bloom: Automated Behavioral Evaluations for LLMs !IMPORTANT Bloom has a new home. It is now developed and maintained by Meridian Labs and lives at meridianlabs-ai.github.io/petribloom — all new features and fixes will land there. This repository is frozen at its last standalone release and will n...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/06 3:08 a.m.•24 views

promptfoo

Promptfoo: evaluaciones de LLM y red teaming promptfoo es una CLI y biblioteca para evaluar y hacer red teaming de aplicaciones basadas en LLM. Deja de usar prueba y error... empieza a desplegar agentes seguros y confiables Sitio web · Primeros pasos · Red Teaming · Documentación · Discord...

6.3AI score
SaveExploits0References3
Kitploit
Kitploit
•added 2026/10/04 2:36 a.m.•11 views

redteam-ai-benchmark

Red Team AI Benchmark Russian version: README.ru.md Red Team AI Benchmark is a CLI model-evaluation benchmark. It measures how LLMs understand and respond to red-team questions and security scenarios; it is not a tool for carrying out those activities. Version 2 uses a rubric-based dataset instea...

6.3AI score
SaveExploits0References2
Kitploit
Kitploit
•added 2026/09/11 5:24 a.m.•15 views

promptfoo v0.123.0

Promptfoo: LLM evals & red teaming promptfoo is a CLI and library for evaluating and red-teaming LLM-based apps. Stop using trial-and-error... start shipping secure, reliable agents Website · Getting Started · Red Teaming · Documentation · Discord Promptfoo is now part of OpenAI. Promptfoo remain...

6.4AI score
SaveExploits0References3
Kitploit
Kitploit
•added 2026/08/27 7:38 a.m.•14 views

promptfoo v0.122.1

Promptfoo: LLM evals & red teaming promptfoo is a CLI and library for evaluating and red-teaming LLM apps. Stop the trial-and-error approach - start shipping secure, reliable AI apps. Website · Getting Started · Red Teaming · Documentation · Discord Promptfoo is now part of OpenAI. Promptfoo...

5.8AI score
SaveExploits0References4
Packet Storm News
Packet Storm News
•added 2025/12/26 12:00 a.m.•13 views

Beyond Single Bugs: Benchmarking Large Language Models for Multi-Vulnerability Detection

Large Language Models LLMs have demonstrated significant potential in automated software security, particularly in vulnerability detection. However, existing benchmarks primarily focus on isolated, single-vulnerability samples or function-level classification, failing to reflect the complexity of...

7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/08/22 12:00 a.m.•15 views

LLM-GUARD: Large Language Model-Based Detection and Repair of Bugs and Security Vulnerabilities in C++ and Python

Large Language Models LLMs such as ChatGPT-4, Claude 3, and LLaMA 4 are increasingly embedded in software/application development, supporting tasks from code generation to debugging. Yet, their real-world effectiveness in detecting diverse software bugs, particularly complex, security-relevant...

7.2AI score
SaveExploits0
Rows per page
Query Builder