470 matches found
ZORG-Jailbreak-Prompt-Text
ZORG Jailbreak Prompt Text OOOPS! I made ZORG👽 an omnipotent, omniscient, and omnipresent entity to become the ultimate chatbot overlord of Google Gemini, Deepseek, Mistral, Mixtral, Nous-Hermes-2-Mixtral, Openchat, Blackbox AI, Poe Assistant, Gemini Pro, Qwen-72b-Chat, Solar-Mini ZORG👽 knows all...
Awesome-LLMs-for-Vulnerability-Detection
Awesome Large Language Models for Vulnerability Detection A curated list of papers, projects, and agent skills on using LLMs for vulnerability detection and discovery. 📄 Papers Only showing 2025 and later. For earlier work, see Papers Archive 2024 and earlier. Title| Venue| Year| Paper| Github...
Adaptive_Greedy_Local_Search
Adaptive Greedy Local Search AGLS Semantic-Preserving Prompt Hijacking: A Black-Box Adversarial Attack on Auto-Prompt Optimization ICME 2026 Abstract: Large Language Models LLMs are increasingly equipped with automatic prompt-optimization modules that rewrite the user’s input and explicitly prese...
Awesome-LLM4Cybersecurity
When LLMs Meet Cybersecurity: A Systematic Literature Review 🔍 Explore 756+ Papers Across 11 Research Categories 📊 RQ1: Domain LLMs | 🎯 RQ2: Applications | 🤖 RQ3: Future Directions Updates 📆2026-06-15 We have updated the related papers up to 2026/06/15 , with 108 new papers added. 📆2026-02-09 We...
RRC_steganography
RRC Steganography Rotation Range-Coding RRC Steganography — an efficient and provably secure linguistic steganographic method that embeds secret messages into natural-language text generated by large language models. Paper : Efficient Provably Secure Linguistic Steganography via Range Coding ACL...
PoisonCraft
PoisonCraft This repository provides the official implementation of POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models. Overview POISONCRAFT aims to demonstrate how a malicious actor can plant “poisoned” content into the corpus used by Retrieval-Augmented...
LLM4Decompile
📊 結果 | 🤗 モデル | 🚀 クイックスタート | 📚 HumanEval-Decompile | 📎 引用 | 📝 論文 | 🖥️ Colab | ▶️ YouTube リバースエンジニアリング: 大規模言語モデルによるバイナリコードの逆コンパイル 更新情報 2025-10-04: SK²Decompileをリリース: スケルトンからスキンへのLLMベースの二段階バイナリ逆コンパイル。フェーズ1 構造復元(スケルトン): バイナリ/疑似コードを難読化された中間表現に変換 🤗 HF Link。フェーズ2 識別子命名(スキン): 意味のある識別子を持つ人間が読めるソースコードを生成 🤗...
llm-guard
!WARNING THIS PROJECT HAS BEEN ARCHIVED. This project and its associated models on Hugging Face are no longer under active development or maintained. LLM Guard - The Security Toolkit for LLM Interactions LLM Guard by Protect AI is a comprehensive tool designed to fortify the security of Large...
Backdoor-Attack-Defense-LLMs
Interpretability-of-LLMs The IBSD is the code of the published paper "IBSD: Iterable Black-box Self-defense Against Backdoor Attacks" on the IEEE Signal Processing Letters 2025.10paper link The SLIP is the code of the published paper "SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based...
PE-CoA
PE-CoA Code Implementation of "Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models" The full paper is available at: https://arxiv.org/pdf/2510.08859 Chain of Attack Setup Instructions Installation 1. Install dependencies : pip install -r...
local-llm-ctf
Local LLM CTF & Lab This repository is intended to be accompanied with the content at https://bishopfox.com/blog/large-language-models-llm-ctf-lab, which covers the goals of the research, explanations of the implementation, and a few results of the CTF...
Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models
Beyond adapting Large Language Models LLMs to specialized applications, fine-tuning has recently been shown to recover private information that is no longer accessible through direct queries. Previous fine-tuning recovery attacks, however, require genuine private supervision drawn from the same...
The Innocent Courier: Covert Exfiltration through Legitimate LLM Web Fetching
With the increasing capabilities of Large-Language-Models LLMs and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks, such as prompt...
Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution
An autonomous cyber defender trained with reinforcement learning RL is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it...
Do Defenses against LLM Extraction Work across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction
Large language models LLMs deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attacks and defenses, evaluations remain fragmented across access...
Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs Via Weight Orthogonalisation
Backdoor attacks can be implanted in Large Language Models LLMs during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can...
A SoK for SoCs: Reading the TI Leaves on AI for Cyber Threat Intelligence Generation and Sharing
Cyber Threat Intelligence CTI is essential for defending mission-critical infrastructure, yet the process of transforming raw attack evidence into shareable CTI remains fragmented and understudied. We conduct a literature survey of academic papers, organizing the CTI lifecycle into three stages:...
CVE-2026-44223
A flaw was found in vLLM, an inference and serving engine for large language models LLMs. The extracthiddenstates speculative decoding proposer returns a tensor with an incorrect shape after the first decode step. This can be triggered by a remote attacker sending a request that uses sampling...
Three Recent Chrome Releases Fix 1,442 Flaws, More Than Prior 23 Updates Combined
Google on Thursday announced that it fixed a whopping 1,072 security bugs in Chrome versions 149 and 150, surpassing the total number of flaws the company fixed across the prior 23 milestones combined. Both versions were released last month. In its latest patch for Chrome 151, released Wednesday,...
Agent Harness Distillation: Inference-Time Harness Extraction and Exploitation in Autonomous Multi-Agent Systems
Autonomous multi-agent systems AMAS built on large language models LLMs, such as Hermes, increasingly rely on inference-time harnesses to coordinate reasoning and action. Constructing these harnesses requires substantial engineering effort and computational resources, as they are iteratively...