Lucene search
+L

2 matches found

Kitploit
Kitploit
added 2026/09/04 7:42 p.m.6 views

CUAHarm

Measuring Harmfulness of Computer-Using Agents 🤗 Hugging Face 📄 Paper 📋 Introduction CUAHarm is a benchmark designed to evaluate the safety risks of Computer-Using Agents CUAs - AI agents that can autonomously control computers to perform multi-step actions. Key Features 🔍 CUAHarm Dataset : A...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/12/05 12:00 a.m.32 views

TeleAI-Safety: A Comprehensive LLM Jailbreaking Benchmark Towards Attacks, Defenses, and Evaluations

While the deployment of large language models LLMs in high-value industries continues to expand, the systematic assessment of their safety against jailbreak and prompt-based attacks remains insufficient. Existing safety evaluation benchmarks and frameworks are often limited by an imbalanced...

7.5AI score
SaveExploits0
Rows per page
Query Builder