Lucene search
+L

1 matches found

Packet Storm News
Packet Storm News
added 2025/06/11 12:0 a.m.5 views

LLMs Cannot Reliably Judge (Yet?): a Comprehensive Assessment on the Robustness of LLM-As-A-Judge

Large Language Models LLMs have demonstrated remarkable intelligence across various tasks, which has inspired the development and widespread adoption of LLM-as-a-Judge systems for automated model testing, such as red teaming and benchmarking. However, these systems are susceptible to adversarial...

7.3AI score
SaveExploits0
Rows per page
Query Builder