Lucene search
+L

2 matches found

Kitploit
Kitploit
added 2026/09/06 2:35 a.m.4 views

redteam-ai-benchmark

Red Team AI Benchmark Russian version: README.ru.md Red Team AI Benchmark is a CLI model-evaluation benchmark. It measures how LLMs understand and respond to red-team questions and security scenarios; it is not a tool for carrying out those activities. Version 2 uses a rubric-based dataset instea...

6AI score
SaveExploits0References2
Packet Storm News
Packet Storm News
added 2025/07/28 12:00 a.m.10 views

Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition

Recent advances have enabled LLM-powered AI agents to autonomously execute complex tasks by combining language model reasoning with tools, memory, and web access. But can these systems be trusted to follow deployment policies in realistic environments, especially under attack? To investigate, we...

7.2AI score
SaveExploits0
Rows per page
Query Builder