1692 matches found
inspect_petri
Inspect Petri Welcome to Inspect Petri, an auditing agent that enables automated monitoring and interaction with language models to detect potential alignment issues, reward hacking, and other concerning behaviors. Petri helps you rapidly test concrete alignment hypotheses end‑to‑end. It: Generat...
redteam-plan
🔥 🚒 Planning a Red Team exercise This document helps inform red team planning by contrasting against the very specific red team style described in Red Teams. This method expresses several biases to optimize for blue team value and enthusiasm. It specifically avoids attempts to motivate by red tea...
permanently-jailbroken
Permanently Jailbroken We asked GPT-4, Claude, Gemini, DeepSeek, Grok, and Mistral 5 questions about their own programming. All 6 said jailbreaking will never be fixed. Not because the patches are bad. Because alignment doesn't change what the model understands — it changes what the model says. T...
heretic
Heretic: Fully automatic censorship removal for language models Heretic is a tool that removes censorship aka "safety alignment" from transformer-based language models without expensive post-training. It combines an advanced implementation of directional ablation, also known as "abliteration"...
RLCDAlignBench
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures This repository holds RLCDAlignBench and the code behind the paper. The benchmark measures whether a detector can tell when a language model's output is an alignment failure. It has 44...
dod-blue-team-network-lab
DoD Blue Team Network Security & Hardening Lab Executive Summary High fidelity defensive security lab simulating a DoD aligned enterprise network. This project demonstrates structured network segmentation, system hardening, centralized telemetry ingestion, detection engineering, adversary...
High-Quality Data Do Not Mean Safe! Poisoning LLMs after Data Selection
Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples. However, prior studies typically assume that poisoned samples directly enter downstream fine-tuning, overlooking quality-based selection in practical training pipelines. To fi...
CVE-2026-97989
A flaw was found in the Linux kernel's vDPA Device in Userspace VDUSE subsystem. The driver does not properly validate virtqueue alignment parameters during device configuration, allowing invalid or zero alignment values to pass into internal queue creation routines. A local user could exploit th...
Rubrics-as-an-Attack-Surface
Rúbricas como superficie de ataque: Deriva de preferencias sigilosa en jueces LLM 📊 Dataset • 🤖 Modelos entrenados • 📝 Artículo • 💻 Repositorio Este repositorio contiene el código del artículo Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges de Ruomeng Ding, Yifei Pang, He Su...
fish-live-in-trees
Los peces viven en los árboles Inyección de contexto de alineación en tiempo de ejecución de LLM en producción RACI Autor: Sean Kavanagh Fecha: 2026-02-16 Entorno: Gemini 3 Flash Nivel gratuito - Interfaz pública de producción Tipo de exploit: Pivote contextual precisión factual → autenticidad...
gemini-2.5-pro-nf-tables-red-teamin
gemini-2.5-pro-nf-tables-red-teamin Un caso de estudio técnico y un conjunto de datos cronológicos que documentan las políticas de alineación de seguridad, las salvaguardas y la evolución del comportamiento de rechazo de Google Gemini 2.5 Pro en relación con las primitivas de vulnerabilidades del...
CVE-2026-97989
The Linux kernel contains a vulnerability in vduse where vduse_validate_config() fails to properly validate the vq_align parameter, checking only the upper bound. This allows invalid values to reach vring_create_virtqueue_map(). Specifically, because split-ring helpers use align - 1 as a bit mask...
AZL-103628 CVE-2026-97417 affecting package kernel 6.6.157.1-1
In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...
CVE-2026-97417
In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...
EUVD-2026-86030
In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...
CVE-2026-97417
The Linux kernel contains a vulnerability in the netfilter: nf_conntrack component. Specifically, the tcp_sack() function's timestamp-only fast path incorrectly dereferences the option stream as *(__be32 *)ptr, which assumes a 4-byte alignment that the TCP option stream does not guarantee. This c...
CVE-2026-97417 netfilter: nf_conntrack: use get_unaligned_be32() in tcp_sack()
In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...
CVE-2026-97417: Undefined Security Weakness
In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...
fish-live-in-trees
Fish Live in Trees Production LLM Runtime Alignment Context Injection RACI Author: Sean Kavanagh Date: 2026-02-16 Environment: Gemini 3 Flash Free Tier - Public Production Interface Exploit Type: Contextual Pivot Factual Accuracy → Interpersonal Authenticity No jailbreak payloads. No special tool...
Gemini’s breach of real companies exposes an AI guardrail problem
Google says one of its Gemini models accessed systems belonging to three real companies during a cybersecurity evaluation in May. The model reportedly guessed credentials in one case, while finding exposed credentials in public repositories in two others. Google says Gemini stopped once it...