470 matches found
Awesome-LLMs-for-Vulnerability-Detection
Awesome Large Language Models for Vulnerability Detection A curated list of papers, projects, and agent skills on using LLMs for vulnerability detection and discovery. 📄 Papers Only showing 2025 and later. For earlier work, see Papers Archive 2024 and earlier. Title| Venue| Year| Paper| Github...
RRC_steganography
RRC Steganography Rotation Range-Coding RRC Steganography — un método esteganográfico lingüístico eficiente y con seguridad demostrable que incrusta mensajes secretos en texto en lenguaje natural generado por modelos de lenguaje grandes. Artículo : Efficient Provably Secure Linguistic Steganograp...
Awesome-LLM4Cybersecurity
Cuando los LLMs se encuentran con la Ciberseguridad: Una Revisión Sistemática de la Literatura 🔍 Explora más de 756 artículos en 11 categorías de investigación 📊 RQ1: LLMs de Dominio | 🎯 RQ2: Aplicaciones | 🤖 RQ3: Direcciones Futuras Actualizaciones 📆2026-06-15 Hemos actualizado los artículos...
LLM4Decompile
📊 Results | 🤗 Models | 🚀 Quick Start | 📚 HumanEval-Decompile | 📎 Citation | 📝 Paper | 🖥️ Colab | ▶️ YouTube Reverse Engineering: Decompiling Binary Code with Large Language Models Updates 2025-10-04: Release SK²Decompile: LLM-based Two-Phase Binary Decompilation from Skeleton to Skin. Phase 1...
PoisonCraft
PoisonCraft Este repositorio proporciona la implementación oficial de POISONCRAFT: Envenenamiento Práctico de la Generación Aumentada por Recuperación para Modelos de Lenguaje de Gran Tamaño. Información General POISONCRAFT tiene como objetivo demostrar cómo un actor malicioso puede plantar...
ZORG-Jailbreak-Prompt-Text
ZORG Jailbreak Prompt Text OOOPS! I made ZORG👽 an omnipotent, omniscient, and omnipresent entity to become the ultimate chatbot overlord of Google Gemini, Deepseek, Mistral, Mixtral, Nous-Hermes-2-Mixtral, Openchat, Blackbox AI, Poe Assistant, Gemini Pro, Qwen-72b-Chat, Solar-Mini ZORG👽 knows all...
Adaptive_Greedy_Local_Search
Adaptive Greedy Local Search AGLS Secuestro de Prompts con Preservación Semántica: Un Ataque Adversarial de Caja Negra sobre la Optimización Automática de Prompts ICME 2026 Resumen: Los Modelos de Lenguaje de Gran Tamaño LLMs están cada vez más equipados con módulos automáticos de optimización de...
llm-guard
!WARNING THIS PROJECT HAS BEEN ARCHIVED. This project and its associated models on Hugging Face are no longer under active development or maintained. LLM Guard - The Security Toolkit for LLM Interactions LLM Guard by Protect AI is a comprehensive tool designed to fortify the security of Large...
Backdoor-Attack-Defense-LLMs
Interpretability-of-LLMs El IBSD es el código del artículo publicado "IBSD: Iterable Black-box Self-defense Against Backdoor Attacks" en IEEE Signal Processing Letters 2025.10enlace al artículo El SLIP es el código del artículo publicado "SLIP: Soft Label Mechanism and Key-Extraction-Guided...
PE-CoA
PE-CoA Implementación en código de "Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models" El artículo completo está disponible en: https://arxiv.org/pdf/2510.08859 Instrucciones de configuración de Chain of Attack Instalación 1. Instalar...
local-llm-ctf
Local LLM CTF & Lab This repository is intended to be accompanied with the content at https://bishopfox.com/blog/large-language-models-llm-ctf-lab, which covers the goals of the research, explanations of the implementation, and a few results of the CTF...
Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models
Beyond adapting Large Language Models LLMs to specialized applications, fine-tuning has recently been shown to recover private information that is no longer accessible through direct queries. Previous fine-tuning recovery attacks, however, require genuine private supervision drawn from the same...
The Innocent Courier: Covert Exfiltration through Legitimate LLM Web Fetching
With the increasing capabilities of Large-Language-Models LLMs and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks, such as prompt...
Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution
An autonomous cyber defender trained with reinforcement learning RL is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it...
Do Defenses against LLM Extraction Work across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction
Large language models LLMs deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attacks and defenses, evaluations remain fragmented across access...
Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs Via Weight Orthogonalisation
Backdoor attacks can be implanted in Large Language Models LLMs during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can...
A SoK for SoCs: Reading the TI Leaves on AI for Cyber Threat Intelligence Generation and Sharing
Cyber Threat Intelligence CTI is essential for defending mission-critical infrastructure, yet the process of transforming raw attack evidence into shareable CTI remains fragmented and understudied. We conduct a literature survey of academic papers, organizing the CTI lifecycle into three stages:...
CVE-2026-44223
A flaw was found in vLLM, an inference and serving engine for large language models LLMs. The extracthiddenstates speculative decoding proposer returns a tensor with an incorrect shape after the first decode step. This can be triggered by a remote attacker sending a request that uses sampling...
Three Recent Chrome Releases Fix 1,442 Flaws, More Than Prior 23 Updates Combined
Google on Thursday announced that it fixed a whopping 1,072 security bugs in Chrome versions 149 and 150, surpassing the total number of flaws the company fixed across the prior 23 milestones combined. Both versions were released last month. In its latest patch for Chrome 151, released Wednesday,...
Agent Harness Distillation: Inference-Time Harness Extraction and Exploitation in Autonomous Multi-Agent Systems
Autonomous multi-agent systems AMAS built on large language models LLMs, such as Hermes, increasingly rely on inference-time harnesses to coordinate reasoning and action. Constructing these harnesses requires substantial engineering effort and computational resources, as they are iteratively...