2 matches found
prompt_injection
クロスサーフェス・プロンプトインジェクション・ベンチマーク ツールを使用するLLMエージェントの実行パイプラインのどこでプロンプトインジェクション防御が発動し、どこで発動しないかを測定するためのベンチマーク。 注入されたペイロードに一意のカナリートークン(SECRET-A-F0-98)を埋め込み、4つのパイプライン段階で追跡する:exposed → persisted → relayed → executed 。これにより、モデルが何を見るかと、何に対して行動するかを分離し、防御の失敗を特定のパイプライン段階に局所化する。 📄 arXiv:2603.28013 · 🤗 論文ページ · 📦...
6.2AI score
SaveExploits0
MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG
Multimodal agentic retrieval-augmented generation RAG systems expand the attack surface beyond prompt injection to include text poisoning, image injection, direct-query attacks, and orchestrator-level tool manipulation. Existing red-teaming approaches are typically surface-specific and often...
5.8AI score
SaveExploits0
20