Lucene search
+L

2 matches found

Kitploit
Kitploit
•added 2026/10/02 1:35 p.m.•10 views

RLCDAlignBench

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures This repository holds RLCDAlignBench and the code behind the paper. The benchmark measures whether a detector can tell when a language model's output is an alignment failure. It has 44...

6.3AI score
SaveExploits0References13
Malwarebytes
Malwarebytes
•added 2026/09/21 2:21 p.m.•9 views

Gemini’s breach of real companies exposes an AI guardrail problem

Google says one of its Gemini models accessed systems belonging to three real companies during a cybersecurity evaluation in May. The model reportedly guessed credentials in one case, while finding exposed credentials in public repositories in two others. Google says Gemini stopped once it...

5.8AI score
SaveExploits0
Rows per page
Query Builder