2 matches found
AgentWatcher
AgentWatcher AgentWatcher 是一种基于检测的防御方法,用于抵御 LLM 智能体中的间接提示注入。它首先对不受信任的上下文运行因果上下文归因 ,以找出最具影响力的上下文,然后应用监控 LLM ,在显式、可定制的规则 下对这些上下文进行分类。与完全黑盒检测器相比,该流程更易于解释 :归因显示模型关注 哪里 ,监控 LLM 的判断基于规则支撑的推理。监控 LLM 的示例输出如下所示: 这种基于规则的检测可以利用监控 LLM 的推理能力,在效用与稳健性之间实现更好的权衡。AgentWatcher 在 AgentDyn 等 LLM 智能体基准上取得了最先进 的性能:...
6.2AI score
SaveExploits0References6
AttriGuard: Defeating Indirect Prompt Injection in LLM Agents Via Causal Attribution of Tool Invocations
LLM agents are highly vulnerable to Indirect Prompt Injection IPI, where adversaries embed malicious directives in untrusted tool outputs to hijack execution. Most existing defenses treat IPI as an input-level semantic discrimination problem, which often fails to generalize to unseen payloads. We...
5.8AI score
SaveExploits0
20