4735 matches found
Adaptive_Greedy_Local_Search
Adaptive Greedy Local Search AGLS 意味保存型プロンプトハイジャック:自動プロンプト最適化に対するブラックボックス敵対的攻撃(ICME 2026) 概要:...
SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Models
Deployment-time safety of language models is commonly implemented through runtime guardrails such as input moderation, routing, retrieval verification, and output filtering. Existing deployment frameworks provide increasingly capable mechanisms for these functions, but offer limited guidance on h...
llm-guard
!WARNING このプロジェクトはアーカイブされました。 このプロジェクトおよび Hugging Face 上の関連モデルは、もはや積極的な開発やメンテナンスは行われていません。 LLM Guard - LLMインタラクションのためのセキュリティツールキット Protect AI による LLM Guard は、大規模言語モデル(LLM)のセキュリティを強化するために設計された包括的なツールです。 ドキュメント | プレイグラウンド | チェンジログ LLM Guardとは?...
Backdoor-Attack-Defense-LLMs
Interpretability-of-LLMs IBSDは、IEEE Signal Processing Letters(2025.10)に掲載された論文「IBSD: Iterable Black-box Self-defense Against Backdoor Attacks」のコードです。論文リンク SLIPは、2026-ACL-findingsに掲載された論文「SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in...
PE-CoA
PE-CoA 「パターン強化型マルチターン・ジェイルブレイキング:大規模言語モデルにおける構造的脆弱性の悪用」のコード実装 論文の全文は以下から入手できます:https://arxiv.org/pdf/2510.08859 Chain of Attackのセットアップ手順 インストール 1. 依存関係をインストール : pip install -r requirements.txt APIキーの設定 以下のAPIキーを config.py と common.py で設定してください: 必須APIキー 1. OpenAI APIキー (最小要件): In config.py...
Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models
Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with...
The Poisoned Conversation: Privacy-Leaking Watermarks in Unified Multimodal Models
Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information...
AutoDP-LLM: Automating Data Pre-Processing for Intrusion Detection Systems Using Large Language Models
The increasing complexity and scale of modern cyber-attacks demand intelligent and computationally efficient Intrusion Detection Systems IDS. However, designing effective data pre-processing pipelines traditionally involves substantial trial-and-error effort and repeated evaluation of alternative...
Forecasting Cybersecurity Incidents Using Geopolitical Data and Large Language Models
Predicting security incidents is a profound task critical for informing proactive defensive measures and cyber-insurance policies. Prior work tackling this problem mainly utilized structured, manually defined features based on network measurements e.g., protocol misconfigurations. Still, despite...
Static Bootstrap Placement for Encrypted Language Model Decoding
Language models increasingly serve prompts that carry private data, and secure inference under homomorphic encryption lets a client outsource the computation without revealing the prompt. Existing secure inference systems run a forward pass without consuming a token under encryption, and generati...
local-llm-ctf
ローカルLLM CTF & ラボ このリポジトリは、https://bishopfox.com/blog/large-language-models-llm-ctf-lab のコンテンツと併せて利用することを意図しています。この記事では、研究の目的、実装の説明、およびCTFのいくつかの結果について説明しています。...
Learning-to-Detect
Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models Official implementation of “Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models.” This repository contains the data-processing, hidden-state extraction, classifier training, safety-pattern...
Picklescan is vulnerable to RCE through missing detection when calling numpy.f2py.crackfortran.myeval
SummaryPicklescan uses numpy.f2py.crackfortran.myeval, which is a function in numpy to execute remote pickle files. DetailsThe attack payload executes in the following steps:- First, the attacker crafts the payload by calling the numpy.f2py.crackfortran.myeval function in its reduce method- Then,...
ODPure
ODPure: アンサンブル汚染コンセンサスによる物体検出のバックドア浄化 物体検出器に特化した、バックドア攻撃に対する初の入力段階ブラックボックス浄化防御。 📖 概要 ODPure は、物体検出器をバックドア攻撃から防御するために特化した、初の入力段階・ブラックボックス浄化フレームワークです。これは、新しい Corruption-Reconstruction-Selection CRS パラダイムを実用化します: 1. 汚染 Corrupts – 多様な摂動で入力画像を汚染し、トリガーパターンを破壊 2. 再構成 Reconstructs –...
Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models
Beyond adapting Large Language Models LLMs to specialized applications, fine-tuning has recently been shown to recover private information that is no longer accessible through direct queries. Previous fine-tuning recovery attacks, however, require genuine private supervision drawn from the same...
The Innocent Courier: Covert Exfiltration through Legitimate LLM Web Fetching
With the increasing capabilities of Large-Language-Models LLMs and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks, such as prompt...
Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution
An autonomous cyber defender trained with reinforcement learning RL is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it...
Do Defenses against LLM Extraction Work across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction
Large language models LLMs deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attacks and defenses, evaluations remain fragmented across access...
Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
As multi-agent systems enter high-stakes domains, the possibility that agents may circumvent safety boundaries is a growing concern. Prior work has examined this risk primarily in adversarial settings, where agents are instructed or rewarded to communicate covertly and evade oversight. We show th...
Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics
Recently, Vision-Language-Action VLA models have revolutionized robotic manipulation by seamlessly integrating visual perception, language understanding, and action generation in an end-to-end learning framework. However, since these models are designed to interact directly with the physical worl...