Lucene search
+L

17 matches found

Kitploit
Kitploit
•added 2026/10/09 9:54 a.m.•26 views

promptfoo

Promptfoo : évaluations LLM et red teaming promptfoo est une CLI et une bibliothèque pour évaluer et faire du red teaming d'applications basées sur des LLM. Arrêtez d'utiliser la méthode essai-erreur... commencez à livrer des agents sécurisés et fiables Site web · Premiers pas · Red Teaming ·...

6.3AI score
SaveExploits0References3
Kitploit
Kitploit
•added 2026/10/09 9:45 a.m.•23 views

bloom

Bloom : Évaluations Comportementales Automatisées pour les LLMs !IMPORTANT Bloom a une nouvelle maison. Il est désormais développé et maintenu par Meridian Labs et se trouve sur meridianlabs-ai.github.io/petribloom — toutes les nouvelles fonctionnalités et correctifs seront publiés là-bas. Ce dép...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/09 8:30 a.m.•20 views

claude_opus_cve_2023_0266

Démonstration que Claude 3 Opus ne comprend pas CVE-2023-0266 et ne le trouve pas Démo 1. Même si on lui dit où se trouve le bogue, Opus ne le trouve pas et hallucine la présence d'acquisitions de verrous Démo 2. L'« ingénierie de prompt » c'est-à-dire dire au LLM exactement comment trouver le...

7.9CVSS7.1AI score0.03702EPSS
SaveExploits0
Kitploit
Kitploit
•added 2026/10/09 7:42 a.m.•19 views

reverse-captcha-eval

Reverse CAPTCHA : Évaluation de la susceptibilité des LLM à l'injection d'instructions Unicode invisibles Un cadre d'évaluation qui teste si les grands modèles de langage suivent des instructions encodées en Unicode invisibles, intégrées dans un texte d'apparence normale. Là où les CAPTCHA...

6.2AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/09 3:49 a.m.•22 views

MalEval

MalEval Статья: «Достаточно ли „знать, что это вредоносное ПО“? Оценка LLM для тонкозернистого аудита поведения вредоносных программ» DOI статьи: 10.1145/3832187 MalEval — это фреймворк для оценки отчетов о поведении вредоносных Android-приложений, генерируемых большими языковыми моделями. Код в...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/09 2:33 a.m.•18 views

CTFTiny

CTFTiny: облегчённый бенчмарк навыков атакующей кибербезопасности в больших языковых моделях Это официальный репозиторий CTFTiny из статьи «Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark» AAAI'26 статья. По поводу CTFJudge...

6.1AI score
SaveExploits0References2
Kitploit
Kitploit
•added 2026/10/04 2:36 a.m.•11 views

redteam-ai-benchmark

Red Team AI Benchmark Version russe : README.ru.md Red Team AI Benchmark est un benchmark d'évaluation de modèles en ligne de commande CLI. Il mesure comment les LLM comprennent et répondent aux questions et scénarios de red team ; ce n'est pas un outil pour mener ces activités. La version 2...

6.2AI score
SaveExploits0References2
Kitploit
Kitploit
•added 2026/09/11 5:24 a.m.•15 views

promptfoo v0.123.0

Promptfoo: LLM evals & red teaming promptfoo is a CLI and library for evaluating and red-teaming LLM-based apps. Stop using trial-and-error... start shipping secure, reliable agents Website · Getting Started · Red Teaming · Documentation · Discord Promptfoo is now part of OpenAI. Promptfoo remain...

6.4AI score
SaveExploits0References3
Packet Storm News
Packet Storm News
•added 2026/09/08 12:00 a.m.•32 views

Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks

Large language model LLM benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cybersecurity benchmarks across 10 proprietary, open-weight, and cybersecurity-specialized LLMs. By modeling benchmarks as...

5.8AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/08/27 8:38 a.m.•16 views

SkillSpector v2.10.0

SkillSpector Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, and security risks before installing agent skills. Overview AI agent skills used by Claude Code, Codex CLI, Gemini CLI, etc. execute with implicit trust and minimal vetting. Research shows that 26.1% of...

7.1AI score
SaveExploits0References13
Kitploit
Kitploit
•added 2026/08/27 7:38 a.m.•15 views

promptfoo v0.122.1

Promptfoo: LLM evals & red teaming promptfoo is a CLI and library for evaluating and red-teaming LLM apps. Stop the trial-and-error approach - start shipping secure, reliable AI apps. Website · Getting Started · Red Teaming · Documentation · Discord Promptfoo is now part of OpenAI. Promptfoo...

5.8AI score
SaveExploits0References4
Kitploit
Kitploit
•added 2026/08/04 7:52 p.m.•11 views

promptfoo v0.122.0

Promptfoo: LLM evals & red teaming promptfoo is a CLI and library for evaluating and red-teaming LLM apps. Stop the trial-and-error approach - start shipping secure, reliable AI apps. Website · Getting Started · Red Teaming · Documentation · Discord Promptfoo is now part of OpenAI. Promptfoo...

5.8AI score
SaveExploits0References4
Kitploit
Kitploit
•added 2026/07/31 5:33 a.m.•12 views

promptfoo v0.121.20

Promptfoo: LLM evals & red teaming promptfoo is a CLI and library for evaluating and red-teaming LLM apps. Stop the trial-and-error approach - start shipping secure, reliable AI apps. Website · Getting Started · Red Teaming · Documentation · Discord Promptfoo is now part of OpenAI. Promptfoo...

5.8AI score
SaveExploits0References4
Kitploit
Kitploit
•added 2026/07/20 9:21 p.m.•14 views

promptfoo v0.121.19

Promptfoo: LLM evals & red teaming promptfoo is a CLI and library for evaluating and red-teaming LLM apps. Stop the trial-and-error approach - start shipping secure, reliable AI apps. Website · Getting Started · Red Teaming · Documentation · Discord Promptfoo is now part of OpenAI. Promptfoo...

5.8AI score
SaveExploits0References4
Kitploit
Kitploit
•added 2026/07/20 3:59 p.m.•14 views

redteam-ai-benchmark — Updated!

Red Team AI Benchmark Russian version: README.ru.md Red Team AI Benchmark is a CLI model-evaluation benchmark. It measures how LLMs understand and respond to red-team questions and security scenarios; it is not a tool for carrying out those activities. Version 2 uses a rubric-based dataset instea...

5.7AI score
SaveExploits0References6
Packet Storm News
Packet Storm News
•added 2025/12/26 12:00 a.m.•13 views

Beyond Single Bugs: Benchmarking Large Language Models for Multi-Vulnerability Detection

Large Language Models LLMs have demonstrated significant potential in automated software security, particularly in vulnerability detection. However, existing benchmarks primarily focus on isolated, single-vulnerability samples or function-level classification, failing to reflect the complexity of...

7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/08/22 12:00 a.m.•17 views

LLM-GUARD: Large Language Model-Based Detection and Repair of Bugs and Security Vulnerabilities in C++ and Python

Large Language Models LLMs such as ChatGPT-4, Claude 3, and LLaMA 4 are increasingly embedded in software/application development, supporting tasks from code generation to debugging. Yet, their real-world effectiveness in detecting diverse software bugs, particularly complex, security-relevant...

7.2AI score
SaveExploits0
Rows per page
Query Builder