1 matches found
SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation
Evaluation of large language models LLMs for safety, security, and privacy SSP relies heavily on static benchmarks, which suffer from score saturation, data contamination, and aggregation artifacts, and fail to capture sensitivity to linguistic variation. As a result, models that perform well on...
5.8AI score
SaveExploits0
20