Lucene search
+L

1 matches found

Packet Storm News
Packet Storm News
•added 2026/08/10 12:00 a.m.•24 views

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the attacks that succeed. Harmful intent is a property of the prompt...

5.5AI score
SaveExploits0
Rows per page
Query Builder