2 matches found
llm-censorship-steering
LLM 审查引导 本仓库包含论文 引导审查:揭示 LLM “思维”控制的表征向量(作者:Hannah Cyberey 和 David Evans)的代码实现。 我们提出了一种从 LLM 内部找到“引导向量”的方法,用于检测和控制模型输出中的审查程度。欢迎查看这篇 博客文章 以快速了解我们的工作。 试试我们的演示: 🐳 使用 DeepSeek-R1-Distill-Qwen-7B 引导 思想抑制 🦙 使用 Llama-3.1-8B-Instruct 引导 拒绝—顺从 注意: 两个演示都需要 Huggingface 账户。该演示托管在 Huggingface 的 ZeroGPU...
5.9AI score
SaveExploits0
Surgical Repair of Insecure Code Generation in LLMs
Large language models write production code, and yet they routinely introduce well-known vulnerabilities. We show that this is not a knowledge deficit: the same models that generate insecure code, correctly identify and explain the vulnerability when asked directly, this is a gap we call the...
5.8AI score
SaveExploits0
20