Lucene search
+L

2 matches found

Kitploit
Kitploit
โ€ขadded 2026/09/25 7:43 p.m.โ€ข13 views

CUAHarm

Measuring Harmfulness of Computer-Using Agents ๐Ÿค— Hugging Face ๐Ÿ“„ Paper ๐Ÿ“‹ Introduction CUAHarm is a benchmark designed to evaluate the safety risks of Computer-Using Agents CUAs - AI agents that can autonomously control computers to perform multi-step actions. Key Features ๐Ÿ” CUAHarm Dataset : A...

6.2AI score
SaveExploits0References5
Packet Storm News
Packet Storm News
โ€ขadded 2025/12/05 12:00 a.m.โ€ข38 views

TeleAI-Safety: A Comprehensive LLM Jailbreaking Benchmark Towards Attacks, Defenses, and Evaluations

While the deployment of large language models LLMs in high-value industries continues to expand, the systematic assessment of their safety against jailbreak and prompt-based attacks remains insufficient. Existing safety evaluation benchmarks and frameworks are often limited by an imbalanced...

7.5AI score
SaveExploits0
Rows per page
Query Builder