2 matches found
CUAHarm
Measuring Harmfulness of Computer-Using Agents ๐ค Hugging Face ๐ Paper ๐ Introduction CUAHarm is a benchmark designed to evaluate the safety risks of Computer-Using Agents CUAs - AI agents that can autonomously control computers to perform multi-step actions. Key Features ๐ CUAHarm Dataset : A...
6.2AI score
SaveExploits0References5
TeleAI-Safety: A Comprehensive LLM Jailbreaking Benchmark Towards Attacks, Defenses, and Evaluations
While the deployment of large language models LLMs in high-value industries continues to expand, the systematic assessment of their safety against jailbreak and prompt-based attacks remains insufficient. Existing safety evaluation benchmarks and frameworks are often limited by an imbalanced...
7.5AI score
SaveExploits0
20