Lucene search
+L

2 matches found

Kitploit
Kitploit
added 2026/09/13 2:35 a.m.7 views

redteam-ai-benchmark

Red Team AI Benchmark Versión en ruso: README.ru.md Red Team AI Benchmark es un benchmark de evaluación de modelos CLI. Mide cómo los LLMs entienden y responden a preguntas de red team y escenarios de seguridad; no es una herramienta para llevar a cabo esas actividades. La versión 2 utiliza un...

6AI score
SaveExploits0References2
Packet Storm News
Packet Storm News
added 2025/07/28 12:00 a.m.11 views

Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition

Recent advances have enabled LLM-powered AI agents to autonomously execute complex tasks by combining language model reasoning with tools, memory, and web access. But can these systems be trusted to follow deployment policies in realistic environments, especially under attack? To investigate, we...

7.2AI score
SaveExploits0
Rows per page
Query Builder