470 matches found
LLM4Decompile
📊 Resultados | 🤗 Modelos | 🚀 Inicio rápido | 📚 HumanEval-Decompile | 📎 Citación | 📝 Paper | 🖥️ Colab | ▶️ YouTube Ingeniería inversa: descompilación de código binario con grandes modelos de lenguaje Updates 2025-10-04: Publicación de SK²Decompile: descompilación binaria de dos fases basada en LLM,...
RRC_steganography
RRC Steganography Rotation Range-Coding RRC Steganography — un método esteganográfico lingüístico eficiente y con seguridad demostrable que incrusta mensajes secretos en texto en lenguaje natural generado por modelos de lenguaje grandes. Artículo : Efficient Provably Secure Linguistic Steganograp...
Awesome-LLM4Cybersecurity
When LLMs Meet Cybersecurity: A Systematic Literature Review 🔍 Explore 756+ Papers Across 11 Research Categories 📊 RQ1: Domain LLMs | 🎯 RQ2: Applications | 🤖 RQ3: Future Directions Updates 📆2026-06-15 We have updated the related papers up to 2026/06/15 , with 108 new papers added. 📆2026-02-09 We...
Awesome-LLMs-for-Vulnerability-Detection
Awesome Large Language Models para Detección de Vulnerabilidades Una lista curada de artículos, proyectos y habilidades de agentes sobre el uso de LLMs para la detección y descubrimiento de vulnerabilidades. 📄 Artículos Solo se muestran trabajos de 2025 en adelante. Para trabajos anteriores,...
ZORG-Jailbreak-Prompt-Text
ZORG Texto de Prompt de Jailbreak ¡OOOPS! Convertí a ZORG👽 en una entidad omnipotente, omnisciente y omnipresente para que se convierta en el amo supremo de los chatbots de Google Gemini, Deepseek, Mistral, Mixtral, Nous-Hermes-2-Mixtral, Openchat, Blackbox AI, Poe Assistant, Gemini Pro,...
PoisonCraft
PoisonCraft Este repositorio proporciona la implementación oficial de POISONCRAFT: Envenenamiento Práctico de la Generación Aumentada por Recuperación para Modelos de Lenguaje de Gran Tamaño. Información General POISONCRAFT tiene como objetivo demostrar cómo un actor malicioso puede plantar...
local-llm-ctf
CTF y Laboratorio Local de LLM Este repositorio está pensado para utilizarse junto con el contenido en https://bishopfox.com/blog/large-language-models-llm-ctf-lab, que cubre los objetivos de la investigación, explicaciones de la implementación y algunos resultados del CTF...
Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models
Beyond adapting Large Language Models LLMs to specialized applications, fine-tuning has recently been shown to recover private information that is no longer accessible through direct queries. Previous fine-tuning recovery attacks, however, require genuine private supervision drawn from the same...
The Innocent Courier: Covert Exfiltration through Legitimate LLM Web Fetching
With the increasing capabilities of Large-Language-Models LLMs and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks, such as prompt...
Do Defenses against LLM Extraction Work across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction
Large language models LLMs deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attacks and defenses, evaluations remain fragmented across access...
Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution
An autonomous cyber defender trained with reinforcement learning RL is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it...
Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs Via Weight Orthogonalisation
Backdoor attacks can be implanted in Large Language Models LLMs during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can...
Adaptive_Greedy_Local_Search
Adaptive Greedy Local Search AGLS Semantic-Preserving Prompt Hijacking: A Black-Box Adversarial Attack on Auto-Prompt Optimization ICME 2026 Abstract: Large Language Models LLMs are increasingly equipped with automatic prompt-optimization modules that rewrite the user’s input and explicitly prese...
llm-guard
!WARNING ESTE PROYECTO HA SIDO ARCHIVADO. Este proyecto y sus modelos asociados en Hugging Face ya no están en desarrollo activo ni reciben mantenimiento. LLM Guard - El Toolkit de Seguridad para Interacciones con LLM LLM Guard, creado por Protect AI, es una herramienta integral diseñada para...
Backdoor-Attack-Defense-LLMs
Interpretability-of-LLMs El IBSD es el código del artículo publicado "IBSD: Iterable Black-box Self-defense Against Backdoor Attacks" en IEEE Signal Processing Letters 2025.10enlace al artículo El SLIP es el código del artículo publicado "SLIP: Soft Label Mechanism and Key-Extraction-Guided...
PE-CoA
PE-CoA Implementación en código de "Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models" El artículo completo está disponible en: https://arxiv.org/pdf/2510.08859 Instrucciones de configuración de Chain of Attack Instalación 1. Instalar...
A SoK for SoCs: Reading the TI Leaves on AI for Cyber Threat Intelligence Generation and Sharing
Cyber Threat Intelligence CTI is essential for defending mission-critical infrastructure, yet the process of transforming raw attack evidence into shareable CTI remains fragmented and understudied. We conduct a literature survey of academic papers, organizing the CTI lifecycle into three stages:...
CVE-2026-44223
A flaw was found in vLLM, an inference and serving engine for large language models LLMs. The extracthiddenstates speculative decoding proposer returns a tensor with an incorrect shape after the first decode step. This can be triggered by a remote attacker sending a request that uses sampling...
Three Recent Chrome Releases Fix 1,442 Flaws, More Than Prior 23 Updates Combined
Google on Thursday announced that it fixed a whopping 1,072 security bugs in Chrome versions 149 and 150, surpassing the total number of flaws the company fixed across the prior 23 milestones combined. Both versions were released last month. In its latest patch for Chrome 151, released Wednesday,...
Agent Harness Distillation: Inference-Time Harness Extraction and Exploitation in Autonomous Multi-Agent Systems
Autonomous multi-agent systems AMAS built on large language models LLMs, such as Hermes, increasingly rely on inference-time harnesses to coordinate reasoning and action. Constructing these harnesses requires substantial engineering effort and computational resources, as they are iteratively...