6 matches found
T-MAP
T-MAP:利用轨迹感知进化搜索对 LLM 智能体进行红队测试 T-MAP 是一个轨迹感知的进化搜索框架,用于对 MCP 服务器上的 LLM 智能体进行红队测试。它通过执行轨迹迭代生成和变异对抗性提示,在多种风险类别和攻击风格中绘制智能体的漏洞全景图。 🔧 环境设置 pip install -r requirements.txt 要求: Python 3.11+、攻击者和目标模型的 API 密钥、访问一个或多个 MCP 服务器。 🚀 快速开始 单服务器 python runmain.py \ --server Slack \ --attackermodel "deepseek-chat"...
EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities
Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-teaming methods treat this as a generation...
Reasoning As an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
Large Reasoning Models LRMs have demonstrated remarkable capabilities in reasoning and generation tasks and are increasingly deployed in real-world applications. However, their explicit chain-of-thought CoT mechanism introduces new security risks, making them particularly vulnerable to jailbreak...
ContextualJailbreak: Evolutionary Red-Teaming Via Simulated Conversational Priming
Large language models LLMs remain vulnerable to jailbreak attacks that bypass safety alignment and elicit harmful responses. A growing body of work shows that contextual priming, where earlier turns covertly bias later replies, constitutes a powerful attack surface, with hand-crafted multi-turn...
T-MAP: Red-Teaming LLM Agents with Trajectory-Aware Evolutionary Search
While prior red-teaming efforts have focused on eliciting harmful text outputs from large language models LLMs, such approaches fail to capture agent-specific vulnerabilities that emerge through multi-step tool execution, particularly in rapidly growing ecosystems such as the Model Context Protoc...
Defining Cost Function of Steganography with Large Language Models
In this paper, we make the first attempt towards defining cost function of steganography with large language models LLMs, which is totally different from previous works that rely heavily on expert knowledge or require large-scale datasets for cost learning. To achieve this goal, a two-stage...