3 matches found
Perturbed-Embedding-Vectors
ランダム埋め込み摂動によるオープンウェイトLLMのジェイルブレイク:実行、ラベル、コード 論文 Jailbreaking Open-Weight LLMs via Random Embedding Perturbations のコードとデータ。 カリフォルニア大学サンタクルーズ校理論計算機科学科の素晴らしい人々によって執筆されました。 この研究には、Scott Sirri、Vaggos Chatziafratis教授、C. Seshadhri教授、そして私自身による重要な貢献が含まれています。 このリポジトリには、論文のために生成されたすべてのモデル応答、各応答の安全 / 不安全 /...
TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors
Current LLM safety benchmarks largely rely on binary metrics, overlooking how models respond to harmful prompts with varying threat implicitness. We introduce TIER, a Threat Implicitness Benchmark for behavioral safety evaluation of LLMs. TIER covers four risk domains and four threat levels, from...
Sockpuppetting: Jailbreaking LLMs without Optimization through Output Prefix Injection
As open-weight large language models LLMs increase in capabilities, safeguarding them against malicious prompts and understanding possible attack vectors becomes ever more important. While automated jailbreaking methods like GCG Zou et al., 2023 remain effective, they often require substantial...