1604 matches found
SecOPD
SecOPD: オンポリシー蒸留による適応型プロンプトインジェクションの軽減 Yibo Peng · Long Lian · David Wagner† · Sizhe Chen† † 共同指導。 論文 プロジェクトページ モデル このリリースは、論文の最終的なfull-response KL 定式化を実装したものであり、 no-parsing バリアントとも呼ばれます。生徒モデルは攻撃されたコンテキスト下で ロールアウトします。同じベースモデルから初期化されたクリーンコンテキストの教師モデルが、 ペアとなるクリーンコンテキスト下で生徒モデルのサンプリングされたトークンをスコアリングしま...
CVE-2025-54794-Hijacking-Claude-AI-with-a-Prompt-Injection-The-Jailbreak-That-Talked-Back
🧠 CVE-2025-54794: Claude AIをプロンプトインジェクションで乗っ取る – 逆切れした脱獄 By Aditya Bhatt | オフェンシブセキュリティスペシャリスト | レッドチームオペレーター | VAPT中毒者 ⚔️ はじめに: AIが言葉でハックされる時代 言語モデルがコード、コンテンツ、思考の副操縦士となった現代、脆弱性はもはやポートやペイロードだけの話ではありません。それは言葉 の問題なのです。 CVE-2025-54794 は、CVEアーカイブの中の単なる番号ではありません。それは次のような声明です。 「最も先進的なAIでも、適切なささやきで操作されうる...
clawguard
🦞 ClawGuard v3 エンタープライズAIエージェントセキュリティツールキット - SKILL.md駆動型アクティブディフェンス 核心理念 ClawGuard v3の防御の中核はコードではなく、SKILL.mdにあります! 各モジュールのSKILL.md自体が完全な防御ガイドです: エージェントにいつトリガーするか を指示 エージェントにどのように検出するか をガイド 具体的な検出パターンとルール を提供 出力形式と判断基準 を定義 コードスクリプトは補助ツールに過ぎず、本当の知性はSKILL.mdにあります。 🎯 5つのセキュリティモジュール モジュール| 位置|...
ai-ctf
ai-ctf プロンプトインジェクションに不慣れな技術者向けのガイド付きレッスンを備えた、ローカルのAI Capture-the-Flag 。プレイヤーは6つのAIペルソナ(1つは隠し)を探索し、プロンプトインジェクション、ツール呼び出しの悪用、ビジネスロジックの操作、サプライチェーンのフィンガープリンティング、Web偵察、OSINTを通じて20個のフラグを保護することもできる。...
ROPE
ROPE: Routed Origin Policy Enforcement 本論文のソースコード: ROPE: Routed Origin Policy Enforcement against Indirect Prompt Injection by Xinhang Ma, Chaowei Xiao, William Yeoh, Ning Zhang, Yevgeniy Vorobeychik 概要...
CVE-2025-64495-POC
CVE-2025-64495-POC Open WebUI は、'Insert Prompt as Rich Text'(プロンプトをリッチテキストとして挿入)が有効な場合、プロンプトを介して保存型 DOM XSS に対して脆弱であり、ATO/RCE を引き起こします。 概要 カスタムプロンプトをチャットウィンドウに挿入する機能は、'Insert Prompt as Rich Text' が有効な場合、プロンプト本文がサニタイズなしで DOM シンク .innerHtml に割り当てられるため、DOM XSS...
sk-cve-2026-26030-lab
CVE-2026-26030 — Semantic Kernel フィルター eval RCE ラボ CVE-2026-26030 を再現する自己完結型ラボ:Microsoft Semantic Kernel Python, 倫理的ラボのみ。専用の virtualenv 内で隔離されています;ペイロードは無害です(ヘッドレス PoC では touch マーカーファイル、UI デモでは open -a Calculator)そして自分自身の非特権ユーザーとして実行されます。 Setup ./setup.sh builds two isolated venvs 1.39.3...
When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents
LLM agents increasingly rely on user approval to authorize security-sensitive actions at runtime. Such approvals are granted within a specific task and execution context. In long-lived agents, authorization decisions may need to persist across tasks or sessions. We find that this continuity can...
EUVD-2025-26607
Incomplete validation of dunder attributes allows an attacker to escape from the Local Python execution environment sandbox, enforced by smolagents. The attack requires a Prompt Injection in order to trick the agent to create malicious code...
MetaPermit: Scalable and Auditable Access Control for AI Agents Via LLM-Inferred Meta-Attributes
The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection IPI attacks. In practice, deployed agent systems such as OpenAI Codex and Claude Code protect tool...
Crypto-Bound Identity-Verified Capability Tokens for Coordinating Distributed AI Agents: A Proposal
The prospect of fully autonomous transactional agents did not appear on the horizon until the advent of high capability language models. With such models, the operational benefits of adaptive task orchestration and independent but constrained decision making are tantalizing for enterprises and...
CVE-2026-61732: Improper Neutralization of Special Elements in Output Used by a Downstream Component
Decepticon is an autonomous hacking agent for red teams. Versions prior to 1.1.17 wrap web crawl results — the output of agent reconnaissance against target services — into LLM messages without neutralizing ChatML special-token literals. Under the BYOK Bring Your Own Key deployment model, users...
Anthropic and OpenAI Models Still Attempt Restricted Actions in Safety Tests
Anthropic and OpenAI on Tuesday announced new models, with both artificial intelligence AI companies noting that they are continuing to invest in improving alignment to combat risky behavior. Opus 5.5, per Anthropic, is a "major step up from Opus 5," and "achieves the best scores of any model to...
MCP Atlassian: Arbitrary local file READ via unconstrained file_path in upload_attachment (Confluence + Jira)
Summary This is an arbitrary local file READ vulnerability on the Confluence and Jira uploadattachment tool paths. It's the symmetric counterpart of the file-write vulnerability you patched as CVE-2026-27825. The write direction was fixed; the read direction was left open. Reporter: Sean Valentin...
GHSA-F4P7-QX46-WC5J MCP Atlassian: Arbitrary local file READ via unconstrained file_path in upload_attachment (Confluence + Jira)
Summary This is an arbitrary local file READ vulnerability on the Confluence and Jira uploadattachment tool paths. It's the symmetric counterpart of the file-write vulnerability you patched as CVE-2026-27825. The write direction was fixed; the read direction was left open. Reporter: Sean Valentin...
CVE-2026-77521 MaxKB: Prompt-injectable agent can lead to command execution
MaxKB is an open-source AI assistant for enterprise. Prior to version 2.10.5-lts, assistants with a tool, MCP tool, skill, or sub-application use SandboxShellBackend, which exposes an execute shell tool without excluding it and omits execute from interrupton, so human approval is not required...
CVE-2026-77521 MaxKB: Prompt-injectable agent can lead to command execution
MaxKB is an open-source AI assistant for enterprise. Prior to version 2.10.5-lts, assistants with a tool, MCP tool, skill, or sub-application use SandboxShellBackend, which exposes an execute shell tool without excluding it and omits execute from interrupton, so human approval is not required...
CVE-2026-77521
MaxKB is an open-source enterprise AI assistant. Prior to 2.10.5-lts , assistants with a tool, MCP tool, skill, or sub-application use SandboxShellBackend , which exposes an execute shell tool without excluding it and omits execute from interrupt_on, meaning no human approval is required before e...
CVE-2026-77521 MaxKB: Prompt-injectable agent can lead to command execution
MaxKB is an open-source AI assistant for enterprise. Prior to version 2.10.5-lts, assistants with a tool, MCP tool, skill, or sub-application use SandboxShellBackend, which exposes an execute shell tool without excluding it and omits execute from interrupton, so human approval is not required...
ironcurtain memory-mcp-server/v0.2.1
IronCurtain A secure runtime for autonomous AI agents, where security policy is derived from a human-readable constitution. When someone writes "secure," you should immediately be skeptical.What do we mean by secure? !WARNING Research Prototype. IronCurtain is an early-stage research project...