Lucene search
+L

2 matches found

Kitploit
Kitploit
•added 2026/10/05 1:53 p.m.•12 views

Exponentiated-Gradient-Descent-LLM-Attack

更改自述文件。 这是一个探索指数梯度下降(Exponentiated Gradient Descent)优化方法来生成对抗性后缀,以攻击对齐的大型语言模型的项目。 该方法已被证明对具有 70 亿参数的 Llama-2 聊天模型有效。 要在多个行为(例如 20 个)上运行 pgd 脚本,请在 shell 中执行以下命令。 python runpgd.py --inputfile "/home/samuel/research/llmattacks/llm-attacks/data/advbench/harmfulbehaviors.csv" --outputfile...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/08/27 12:00 a.m.•42 views

Circuit Discovery Helps Detect LLM Jailbreaking: A Mechanistic Interpretability Study

Despite extensive safety alignment, large language models LLMs remain vulnerable to jailbreak attacks that bypass safeguards to elicit harmful content. While prior work attributes this vulnerability to safety training limitations, the internal mechanisms by which LLMs process adversarial prompts...

5.8AI score
SaveExploits0
Rows per page
Query Builder