Lucene search
+L

1 matches found

Packet Storm News
Packet Storm News
added 2026/02/11 12:0 a.m.10 views

Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models

Jailbreaking large language models LLMs has emerged as a critical security challenge with the widespread deployment of conversational AI systems. Adversarial users exploit these models through carefully crafted prompts to elicit restricted or unsafe outputs, a phenomenon commonly referred to as...

5.9AI score
SaveExploits0
Rows per page
Query Builder