2 matches found
RLCDAlignBench
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures This repository holds RLCDAlignBench and the code behind the paper. The benchmark measures whether a detector can tell when a language model's output is an alignment failure. It has 44...
6.3AI score
SaveExploits0References13
Gemini’s breach of real companies exposes an AI guardrail problem
Google says one of its Gemini models accessed systems belonging to three real companies during a cybersecurity evaluation in May. The model reportedly guessed credentials in one case, while finding exposed credentials in public repositories in two others. Google says Gemini stopped once it...
5.8AI score
SaveExploits0
20