Lucene search
+L

3 matches found

Kitploit
Kitploit
•added 2026/09/25 7:32 a.m.•4 views

FinRED-paper

FinRED: Financial Red-Teaming Evaluation Dataset A red-team benchmark generation pipeline for safety evaluation in the financial domain. Supplementary Documentation Detailed materials referenced in the paper: docs/expertvalidation.md — Per-question Focus Group Interview FGI results from the 12 FS...

6.1AI score
SaveExploits0References4
Kitploit
Kitploit
•added 2026/09/20 2:35 a.m.•9 views

redteam-ai-benchmark

Red Team AI Benchmark Versión en ruso: README.ru.md Red Team AI Benchmark es un benchmark de evaluación de modelos CLI. Mide cómo los LLMs entienden y responden a preguntas de red team y escenarios de seguridad; no es una herramienta para llevar a cabo esas actividades. La versión 2 utiliza un...

6.1AI score
SaveExploits0References2
Packet Storm News
Packet Storm News
•added 2025/07/28 12:00 a.m.•12 views

Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition

Recent advances have enabled LLM-powered AI agents to autonomously execute complex tasks by combining language model reasoning with tools, memory, and web access. But can these systems be trusted to follow deployment policies in realistic environments, especially under attack? To investigate, we...

7.2AI score
SaveExploits0
Rows per page
Query Builder