Lucene search
+L

4 matches found

Packet Storm News
Packet Storm News
•added 2026/08/20 12:00 a.m.•14 views

AiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-Offs in LLM Safety, Security, and Privacy

The critical failure modes in deployed large language models LLMs are cross-dimensional: a model can score 99.3 in safety alignment while refusing one in three benign queries, or improve across every capability metric while losing 21 points in privacy. Existing evaluation frameworks that assess...

5.6AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/04/23 12:00 a.m.•17 views

AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models

Automated red-teaming methods for large language models typically optimize attack prompts within a fixed, human-designed strategy, leaving the attack strategy itself unchanged. We instead optimize the strategy. We propose AutoRISE, a method that searches over executable attack programs rather tha...

5.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/04/05 12:00 a.m.•16 views

Towards Unveiling Vulnerabilities of Large Reasoning Models in Machine Unlearning

Large language models LLMs possess strong semantic understanding, driving significant progress in data mining applications. This is further enhanced by large reasoning models LRMs, which provide explicit multi-step reasoning traces. On the other hand, the growing need for the right to be forgotte...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/22 12:00 a.m.•14 views

Pushing the Limits of Safety: a Technical Report on the ATLAS Challenge 2025

Multimodal Large Language Models MLLMs have enabled transformative advancements across diverse applications but remain susceptible to safety threats, especially jailbreak attacks that induce harmful outputs. To systematically evaluate and improve their safety, we organized the Adversarial Testing...

7.6AI score
SaveExploits0
Rows per page
Query Builder