Lucene search
+L

4 matches found

Kitploit
Kitploit
added 2026/09/10 8:11 p.m.8 views

BrokenHill

!\ Broken Hill 标志 \https://assets.kitploit.com/production/public/readmes/44608/c14154675c6f33f5c3887887afcdb39351c82f269ee2772cf31847106a20a871.jpg Broken Hill Broken Hill 是一款已产品化、开箱即用的自动化攻击工具,它使用贪婪坐标梯度(greedy coordinate gradient, GCG)攻击生成精心构造的提示词,以绕过大型语言模型(LLM)中的限制。该攻击方法详见 Andy Zou、Zifan...

5.8AI score
SaveExploits0References5
Packet Storm News
Packet Storm News
added 2026/05/27 12:00 a.m.57 views

Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on residual stream activations at each transformer block. We find that refusal is linearly decodable well before the final layer, indicating that...

5.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/11/20 12:00 a.m.49 views

Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models

The growing misuse of Vision-Language Models VLMs has led providers to deploy multiple safeguards, including alignment tuning, system prompts, and content moderation. However, the real-world robustness of these defenses against adversarial attacks remains underexplored. We introduce Multi-Faceted...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/10 12:00 a.m.9 views

DAVSP: Safety Alignment for Large Vision-Language Models Via Deep Aligned Visual Safety Prompt

Large Vision-Language Models LVLMs have achieved impressive progress across various applications but remain vulnerable to malicious queries that exploit the visual modality. Existing alignment approaches typically fail to resist malicious queries while preserving utility on benign ones effectivel...

7.5AI score
SaveExploits0
Rows per page
Query Builder