2 matches found
CUAHarm
Measuring Harmfulness of Computer-Using Agents 🤗 Hugging Face 📄 Paper 📋 Introduction CUAHarm is a benchmark designed to evaluate the safety risks of Computer-Using Agents CUAs - AI agents that can autonomously control computers to perform multi-step actions. Key Features 🔍 CUAHarm Dataset : A...
5.9AI score
SaveExploits0
TeleAI-Safety: A Comprehensive LLM Jailbreaking Benchmark Towards Attacks, Defenses, and Evaluations
While the deployment of large language models LLMs in high-value industries continues to expand, the systematic assessment of their safety against jailbreak and prompt-based attacks remains insufficient. Existing safety evaluation benchmarks and frameworks are often limited by an imbalanced...
7.5AI score
SaveExploits0
20