Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2026/04/06 12:0 a.m.10 views

SALLIE: Safeguarding against Latent Language and Image Exploits

Large Language Models LLMs and Vision-Language Models VLMs remain highly vulnerable to textual and visual jailbreaks, as well as prompt injections arXiv:2307.15043, Greshake et al., 2023, arXiv:2306.13213. Existing defenses often degrade performance through complex input transformations or treat...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/03/15 12:0 a.m.5 views

Activation Surgery: Jailbreaking White-Box LLMs without Touching the Prompt

Most jailbreak techniques for Large Language Models LLMs primarily rely on prompt modifications, including paraphrasing, obfuscation, or conversational strategies. Meanwhile, abliteration techniques also known as targeted ablations of internal components have been used to study and explain LLM...

5.9AI score
SaveExploits0
Rows per page
Query Builder