Lucene search
+L

1 matches found

Packet Storm News
Packet Storm News
added 2025/09/18 12:0 a.m.7 views

Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism Via Probabilistically Ablating Refusal Direction

Jailbreak attacks pose persistent threats to large language models LLMs. Current safety alignment methods have attempted to address these issues, but they experience two significant limitations: insufficient safety alignment depth and unrobust internal defense mechanisms. These limitations make...

7.3AI score
SaveExploits0
Rows per page
Query Builder