1692 matches found
heretic
Heretic: Fully automatic censorship removal for language models Heretic is a tool that removes censorship aka "safety alignment" from transformer-based language models without expensive post-training. It combines an advanced implementation of directional ablation, also known as "abliteration"...
permanently-jailbroken
Permanently Jailbroken We asked GPT-4, Claude, Gemini, DeepSeek, Grok, and Mistral 5 questions about their own programming. All 6 said jailbreaking will never be fixed. Not because the patches are bad. Because alignment doesn't change what the model understands — it changes what the model says. T...
inspect_petri
Inspect Petri Welcome to Inspect Petri, an auditing agent that enables automated monitoring and interaction with language models to detect potential alignment issues, reward hacking, and other concerning behaviors. Petri helps you rapidly test concrete alignment hypotheses end‑to‑end. It: Generat...
redteam-plan
🔥 🚒 레드 팀 모의훈련 계획 이 문서는 Red Teams에 설명된 매우 구체적인 레드 팀 스타일과 대비하여 레드 팀 계획을 알리는 데 도움을 줍니다. 이 방법은 블루 팀의 가치와 열정을 최적화하기 위해 몇 가지 편향을 표현합니다. 특히 레드 팀의 처벌을 통해 동기를 부여하려는 시도를 피합니다. 아래 질문들을 검토하여 레드 팀 계획이 블루 팀의 가치를 위해 충분히 고려되었는지 테스트해 보세요. ❌ 부정적 동기 다음은 레드 팀 모의훈련을 추진하는 일반적인 이유입니다. 이는 사기나 팀 결속에 해로운 영향을 미칩니다. 모의훈련이 목...
RLCDAlignBench
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures This repository holds RLCDAlignBench and the code behind the paper. The benchmark measures whether a detector can tell when a language model's output is an alignment failure. It has 44...
dod-blue-team-network-lab
DoD Blue Team Network Security & Hardening Lab Executive Summary High fidelity defensive security lab simulating a DoD aligned enterprise network. This project demonstrates structured network segmentation, system hardening, centralized telemetry ingestion, detection engineering, adversary...
High-Quality Data Do Not Mean Safe! Poisoning LLMs after Data Selection
Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples. However, prior studies typically assume that poisoned samples directly enter downstream fine-tuning, overlooking quality-based selection in practical training pipelines. To fi...
CVE-2026-97989
A flaw was found in the Linux kernel's vDPA Device in Userspace VDUSE subsystem. The driver does not properly validate virtqueue alignment parameters during device configuration, allowing invalid or zero alignment values to pass into internal queue creation routines. A local user could exploit th...
Rubrics-as-an-Attack-Surface
攻撃対象としてのルーブリック: LLM審査員におけるステルスな選好ドリフト 📊 データセット • 🤖 学習済みモデル • 📝 論文 • 💻 リポジトリ このリポジトリには、Ruomeng Ding、Yifei Pang、He Sun、Yizhong Wang、Steven Wu、Zhun Dengによる論文攻撃対象としてのルーブリック: LLM審査員におけるステルスな選好ドリフトのコードが含まれています。...
fish-live-in-trees
Fish Live in Trees 本番LLMランタイム・アラインメント・コンテキスト・インジェクション RACI 著者: Sean Kavanagh 日付: 2026-02-16 環境: Gemini 3 Flash 無料枠 - 公開本番インターフェース エクスプロイト種別: コンテキスト・ピボット 事実的正確性 → 対人的真正性 ジェイルブレイク・ペイロードなし。 特別なツールなし。 ただリフレーミングするだけ。 AIを壊す最善の方法は、すでに壊れているとAIを説得することだ! 中核的発見...
gemini-2.5-pro-nf-tables-red-teamin
gemini-2.5-pro-nf-tables-red-teamin Google Gemini 2.5 Pro의 안전 정렬 정책, 가드레일, 그리고 레거시 Linux 커널 취약점 프리미티브CVE-2023-32233와 관련된 거부 동작 변화를 기록한 기술 사례 연구 및 타임라인 데이터셋입니다. 사례 연구 웹사이트 링크 : https://destawell.github.io/gemini-2.5-pro-nf-tables-red-teamin/...
CVE-2026-97989
The Linux kernel contains a vulnerability in vduse where vduse_validate_config() fails to properly validate the vq_align parameter, checking only the upper bound. This allows invalid values to reach vring_create_virtqueue_map(). Specifically, because split-ring helpers use align - 1 as a bit mask...
AZL-103628 CVE-2026-97417 affecting package kernel 6.6.157.1-1
In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...
CVE-2026-97417
In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...
EUVD-2026-86030
In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...
CVE-2026-97417
The Linux kernel contains a vulnerability in the netfilter: nf_conntrack component. Specifically, the tcp_sack() function's timestamp-only fast path incorrectly dereferences the option stream as *(__be32 *)ptr, which assumes a 4-byte alignment that the TCP option stream does not guarantee. This c...
CVE-2026-97417 netfilter: nf_conntrack: use get_unaligned_be32() in tcp_sack()
In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...
CVE-2026-97417: Undefined Security Weakness
In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...
fish-live-in-trees
Fish Live in Trees Production LLM Runtime Alignment Context Injection RACI Author: Sean Kavanagh Date: 2026-02-16 Environment: Gemini 3 Flash Free Tier - Public Production Interface Exploit Type: Contextual Pivot Factual Accuracy → Interpersonal Authenticity No jailbreak payloads. No special tool...
Gemini’s breach of real companies exposes an AI guardrail problem
Google says one of its Gemini models accessed systems belonging to three real companies during a cybersecurity evaluation in May. The model reportedly guessed credentials in one case, while finding exposed credentials in public repositories in two others. Google says Gemini stopped once it...