Lucene search
+L

1218 matches found

Kitploit
Kitploit
•added 2026/10/04 6:11 a.m.•12 views

CTFTiny

CTFTiny: Lite Benchmarking Offensive Cyber Skills in Large Language Models This is the official repository for CTFTiny from "Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark" AAAI'26 paper. For CTFJudge, please refer to CTFJud...

6.1AI score
SaveExploits0References2
Kitploit
Kitploit
•added 2026/10/04 6:10 a.m.•28 views

xalgorix

Xalgorix — Open-source AI pentester that proves vulnerabilities Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported — so you get proof, not a pile of maybes to triage...

6AI score
SaveExploits0References3
Kitploit
Kitploit
•added 2026/10/04 6:06 a.m.•20 views

llm-circuit-finder

llm-circuit-finder I replicated Ng's RYS method and found that duplicating 3 specific layers in Qwen2.5-32B boosts reasoning by 17% and duplicating layers 12-14 in Devstral-24B improves logical deduction from 0.22→0.76 on BBH — no training, no weight changes, just routing hidden states through th...

6.4AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/04 5:57 a.m.•20 views

DVAP

DVAP - Damn Vulnerable AI Platform Train. Break. Defend. AI Systems. An open-source platform for AI security training, red/blue teaming, CTF, benchmarking, and research. Runs 100% locally. No cloud, no paid APIs, no data leaves your machine. Prerequisites Docker 24+ and Docker Compose v2 16 GB RA...

6.2AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/10/04 5:39 a.m.•19 views

MalEval

MalEval Article: Is “Knowing It’s Malicious” Enough? Evaluating LLMs for Fine-Grained Malware Behavior Auditing Article DOI: 10.1145/3832187 MalEval is a framework for evaluating Android malware behavior reports generated by large language models. The code in this repository implements two...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/04 5:28 a.m.•12 views

redeval

RedEval - LLM Safety Evaluation Framework A comprehensive framework for evaluating the safety of Large Language Models LLMs through systematic attack and refusal testing. RedEval provides a unified, secure, and extensible platform for assessing LLM robustness against adversarial prompts and harmf...

6.3AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/10/04 5:23 a.m.•23 views

ai-kill-chain

Extended Cyber Kill Chain for AI-Era Threats An update to the Lockheed Martin Cyber Kill Chain for defenders working against LLM and agentic AI attacks. Adds a pre-attack stage for model supply chain compromise. Adds AI-specific sub-techniques to each of the original seven stages. Splits the...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/04 5:15 a.m.•16 views

www-project-top-10-for-large-language-model-applications

www-project-top-10-for-large-language-model-applications OWASP Foundation Web Repository OWASP Top 10 for Large Language Model Applications Welcome to the official repository for the OWASP Top 10 for Large Language Model Applications! About This Repository This repository contains the OWASP Top 1...

5.9AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/04 5:12 a.m.•19 views

windbg-decompile-ext

Windbg Decompile Extension via LLM This project is a Windows x64 WinDbg extension skeleton that resolves a function by name or address, reconstructs a deterministic control-flow view, and asks an LLM directly from the extension to produce pseudocode. Layout src/extension: WinDbg extension DLL and...

6.3AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/10/04 5:01 a.m.•18 views

claude_opus_cve_2023_0266

Demonstration that Claude 3 Opus does not understand CVE-2023-0266 and does not find it Demo 1. Even if told where the bug is Opus does not find it, and hallucinates the presence of lock acquisitions Demo 2. "Prompt engineering" aka telling the LLM exactly how to find the bug also doesn't work De...

7.9CVSS6.9AI score0.03702EPSS
SaveExploits0
Kitploit
Kitploit
•added 2026/10/04 5:00 a.m.•19 views

supply-chain-monitor

Supply Chain Monitor Automated monitoring of the top PyPI and npm packages for supply chain compromise. Polls both registries for new releases, diffs each release against its predecessor, and uses an LLM via Cursor Agent CLI to classify diffs as benign or malicious. Malicious findings trigger a...

6.2AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/04 4:58 a.m.•13 views

quill-router

TrustedRouter End-to-end encrypted LLMs. One API. Privacy with proof. Stop worrying about who can see your prompts. Tell your coding agent to move your project over, pick how private you want to be, pick a model, drop in a key — done. Same API, 30+ models, one key. The gateway runs in hardware...

6.3AI score
SaveExploits0References6
Kitploit
Kitploit
•added 2026/10/04 4:57 a.m.•14 views

ClawGuard

ClawGuard 🛡️ 中文版 Our Project:https://github.com/SafeAgent-Beihang/clawguard ClawGuard is a security toolkit designed to mitigate risks associated with autonomous agents, such as OpenClaw and other LLM-driven entities. As agents gain more autonomy to execute code, access APIs, and manage files,...

6.2AI score
SaveExploits0References7
Kitploit
Kitploit
•added 2026/10/04 4:42 a.m.•16 views

dataset

🚀 CySecBench: Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models 🛡️ The largest and most comprehensive Generative AI-based CyberSecurity-focused Dataset for Benchmarking Large Language Models 🌟 Overview The CySecBench paper offers: 🎯 A cutting-edge...

6.3AI score
SaveExploits0References12
Kitploit
Kitploit
•added 2026/10/04 4:39 a.m.•15 views

reasongate

ReasonGate A self-hostable gate that inspects the text going into and out of an LLM and returns an explainable allow / flag / block decision with a machine-readable audit record for every call. What this is The open-source core is rule-based. It does four things: recognizes known prompt-injection...

6AI score
SaveExploits0References2
Kitploit
Kitploit
•added 2026/10/04 4:36 a.m.•21 views

threat-modeling-mcp-server

Threat Modeling MCP Server A Model Context Protocol MCP server for comprehensive threat modeling with guided code validation. Table of Contents Overview Quick Start Prompts Key Features Prerequisites Installation Running with Kiro CLI Output File Management Quick Reference Tools Overview Threat...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/04 4:28 a.m.•12 views

pike-agent

Pike Agent pike-agent records and analyzes how programs behave on Linux. It traces a program's activity, indexes it into a database, and lets you chat with an LLM agent about it in a TUI. Example of prompts: Crash diagnosis: This program crashed with a bus error. What happened? Race condition...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/04 4:27 a.m.•16 views

reverse-captcha-eval

Reverse CAPTCHA: Evaluating LLM Susceptibility to Invisible Unicode Instruction Injection An evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in otherwise normal-looking text. Where traditional CAPTCHAs exploit tasks humans can...

6.5AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/04 4:17 a.m.•20 views

stride-gpt

STRIDE GPT is an AI-powered threat modelling tool that leverages Large Language Models LLMs to generate threat models and attack trees for a given application based on the STRIDE methodology. Users provide application details, such as the application type, authentication methods, and whether the...

6.1AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/04 4:07 a.m.•9 views

saffron

Saffron-1: Inference Scaling for LLM Safety Assurance 📖 Paper 🛠️ Dependencies The code was tested under the following dependencies: Python 3.12.3 CUDA 12.2 typingextensions==4.14.0 numpy==2.2.6 torch==2.5.1 huggingfacehub==0.30.2 accelerate==1.1.1 datasets==3.1.0 evaluate==0.4.3...

6AI score
SaveExploits0
Rows per page
Query Builder