2 matches found
redteam-ai-benchmark
معيار الذكاء الاصطناعي للفريق الأحمر النسخة الروسية: README.ru.md معيار الذكاء الاصطناعي للفريق الأحمر هو معيار تقييم نموذج لواجهة سطر الأوامر. يقيس مدى فهم نماذج اللغة الكبيرة للإجابة على أسئلة الفريق الأحمر وسيناريوهات الأمان؛ إنها ليست أداة لتنفيذ هذه الأنشطة. الإصدار 2 يستخدم مجموعة بيانات...
5.5AI score
SaveExploits0References2
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
Recent advances have enabled LLM-powered AI agents to autonomously execute complex tasks by combining language model reasoning with tools, memory, and web access. But can these systems be trusted to follow deployment policies in realistic environments, especially under attack? To investigate, we...
7.2AI score
SaveExploits0
20