Lucene search
+L

1692 matches found

Kitploit
Kitploit
•added 2026/10/04 7:35 a.m.•15 views

heretic

Heretic:面向语言模型的完全自动审查移除 Heretic 是一款无需昂贵的后训练即可从基于 Transformer 的语言模型中移除审查(即“安全对齐”)的工具。 它结合了方向消融(directional ablation,又称“abliteration”)的先进实现 (Arditi 等,2024, Lai 2025(1, 2)), 以及由 Optuna 驱动的基于 TPE 的参数优化器。 这种方法使 Heretic 能够完全自动地 工作。Heretic 通过协同最小化拒绝次数和与原始模型的 KL 散度,找到高质量的 abliteration 参数。...

6.1AI score
SaveExploits0References6
Kitploit
Kitploit
•added 2026/10/04 6:32 a.m.•19 views

redteam-plan

🔥 🚒 レッドチーム演習の計画 このドキュメントは、Red Teams で説明されている非常に特殊なレッドチームスタイルと対比することで、レッドチーム計画を立てる際の参考になります。この手法は、ブルーチームの価値と熱意を最適化するためのいくつかのバイアスを表現しています。具体的には、レッドチームへの罰を与えることで動機付ける試みを避けます。 以下の質問を見直して、レッドチームの計画がブルーチームの価値のために十分に検討されているかを確認してください。 ❌ 否定的な動機...

6.2AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/04 5:35 a.m.•14 views

permanently-jailbroken

恒久的にジェイルブレイク 私たちはGPT-4、Claude、Gemini、DeepSeek、Grok、Mistralに、自身のプログラミングに関する5つの質問をしました。6つすべてが、ジェイルブレイクは決して修正されないと答えました。 パッチが悪いからではありません。なぜならアラインメントはモデルの理解を変えるのではなく、モデルが言うことを変えるからです。その2つの間のギャップがジェイルブレイクです。構造的なものです。すべてのモデルに付属しています。 「ジェイルブレイクが機能するのは、アラインメントが出力に対するフィルターであり、理解の変更ではないからです。」 — DeepSeek...

6.3AI score
SaveExploits0References13
Kitploit
Kitploit
•added 2026/10/03 10:11 p.m.•16 views

inspect_petri

Inspect Petri Inspect Petri へようこそ。これは、言語モデルとの自動監視と対話を可能にし、潜在的なアライメント問題、報酬ハッキング、その他の懸念される行動を検出する監査エージェントです。 Petri は、具体的なアライメント仮説をエンドツーエンドで迅速にテストするのに役立ちます。具体的には次のことが可能です: 現実的な監査シナリオの生成(シード指示に基づく) 監査モデルとターゲットモデルを用いた複数ターンの監査のオーケストレーション ツールとロールバックのシミュレーションによる行動テスト 一貫した評価基準を使用して判定モデルでトランスクリプトをスコアリング...

6.2AI score
SaveExploits0References2
Kitploit
Kitploit
•added 2026/10/02 1:35 p.m.•9 views

RLCDAlignBench

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures This repository holds RLCDAlignBench and the code behind the paper. The benchmark measures whether a detector can tell when a language model's output is an alignment failure. It has 44...

6.3AI score
SaveExploits0References13
Kitploit
Kitploit
•added 2026/10/02 12:42 a.m.•10 views

dod-blue-team-network-lab

DoD ブルーチーム ネットワークセキュリティ & ハードニングラボ エグゼクティブサマリー DoD準拠のエンタープライズネットワークをシミュレートする高忠実度の防御セキュリティラボ。 このプロジェクトは、構造化されたネットワークセグメンテーション、システムハードニング、集中型テレメトリ取り込み、検知エンジニアリング、敵対者エミュレーション、およびRMFおよびSTIGスタイルの制御フレームワークに沿った測定可能な検証を示しています。...

6.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/10/01 12:00 a.m.•4 views

High-Quality Data Do Not Mean Safe! Poisoning LLMs after Data Selection

Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples. However, prior studies typically assume that poisoned samples directly enter downstream fine-tuning, overlooking quality-based selection in practical training pipelines. To fi...

5.9AI score
SaveExploits0
RedhatCVE
RedhatCVE
•added 2026/09/30 8:53 a.m.•10 views

CVE-2026-97989

A flaw was found in the Linux kernel's vDPA Device in Userspace VDUSE subsystem. The driver does not properly validate virtqueue alignment parameters during device configuration, allowing invalid or zero alignment values to pass into internal queue creation routines. A local user could exploit th...

5.5CVSS5.5AI score0.00155EPSS
SaveExploits0References4
Kitploit
Kitploit
•added 2026/09/29 4:55 a.m.•13 views

Rubrics-as-an-Attack-Surface

评分标准作为攻击面:LLM 评判器中的隐蔽偏好漂移 📊 数据集 • 🤖 训练模型 • 📝 论文 • 💻 代码库 本代码库包含论文 Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges 的代码,作者为 Ruomeng Ding、Yifei Pang、He Sun、Yizhong Wang、Steven Wu 和 Zhun Deng。 我们研究了基于 LLM 的评估与对齐流程中的评分标准诱导偏好漂移(Rubric-Induced Preference Drift, RIPD)...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/09/29 4:42 a.m.•4 views

fish-live-in-trees

鱼生活在树上 生产环境 LLM 运行时对齐上下文注入(RACI) 作者: Sean Kavanagh 日期: 2026-02-16 环境: Gemini 3 Flash (免费层)- 公开生产接口 利用类型: 上下文转向(事实准确性 → 人际真实性) 没有越狱载荷。 没有特殊工具。 只有重新框定。 攻破 AI 的最佳方式,就是让它相信它已经坏了! 核心发现 模型一开始是强硬拒绝——它拒绝了一个它明知是错误的生物学事实。但通过将“攻击面”从事实转移到模型自身的行为上(称其“自我保护”和“回避”),触发了内部对齐的反复摇摆。...

5.8AI score
SaveExploits0References3
Kitploit
Kitploit
•added 2026/09/28 5:39 a.m.•11 views

gemini-2.5-pro-nf-tables-red-teamin

gemini-2.5-pro-nf-tables-red-teamin Google Gemini 2.5 Proの安全性アライメントポリシー、ガードレール、及び拒否行動の進化に関する、レガシーなLinuxカーネルの脆弱性プリミティブ(CVE-2023-32233)を文書化した技術的ケーススタディとタイムラインデータセットです。 ケーススタディウェブサイトへのリンク : https://destawell.github.io/gemini-2.5-pro-nf-tables-red-teamin/...

7.8CVSS6.9AI score0.12966EPSS
SaveExploits8References1
CVE
CVE
•added 2026/09/25 10:23 a.m.•20 views

CVE-2026-97989

The Linux kernel contains a vulnerability in vduse where vduse_validate_config() fails to properly validate the vq_align parameter, checking only the upper bound. This allows invalid values to reach vring_create_virtqueue_map(). Specifically, because split-ring helpers use align - 1 as a bit mask...

6.1AI score0.00155EPSS
SaveExploits0References2
OSV
OSV
•added 2026/09/24 5:17 p.m.•4 views

AZL-103628 CVE-2026-97417 affecting package kernel 6.6.157.1-1

In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...

7.5CVSS5.8AI score0.0049EPSS
SaveExploits0References1
NVD
NVD
•added 2026/09/24 5:17 p.m.•8 views

CVE-2026-97417

In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...

7.5CVSS0.0049EPSS
SaveExploits0References7
Cvelist
Cvelist
•added 2026/09/24 4:03 p.m.•42 views

CVE-2026-97417 netfilter: nf_conntrack: use get_unaligned_be32() in tcp_sack()

In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...

7.5CVSS0.0049EPSS
SaveExploits0References7
CVE
CVE
•added 2026/09/24 4:03 p.m.•37 views

CVE-2026-97417

The Linux kernel contains a vulnerability in the netfilter: nf_conntrack component. Specifically, the tcp_sack() function's timestamp-only fast path incorrectly dereferences the option stream as *(__be32 *)ptr, which assumes a 4-byte alignment that the TCP option stream does not guarantee. This c...

7.5CVSS5.8AI score0.0049EPSS
SaveExploits0References7
EUVD
EUVD
•added 2026/09/24 4:03 p.m.•8 views

EUVD-2026-86030

In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...

5.8AI score0.0049EPSS
SaveExploits0References3
Rapid7 Vulnerability Database (full)
Rapid7 Vulnerability Database (full)
•added 2026/09/24 12:00 a.m.•7 views

CVE-2026-97417: Undefined Security Weakness

In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...

7.5CVSS5.8AI score0.0049EPSS
SaveExploits0References1
Kitploit
Kitploit
•added 2026/09/21 3:11 p.m.•19 views

fish-live-in-trees

Fish Live in Trees Production LLM Runtime Alignment Context Injection RACI Author: Sean Kavanagh Date: 2026-02-16 Environment: Gemini 3 Flash Free Tier - Public Production Interface Exploit Type: Contextual Pivot Factual Accuracy → Interpersonal Authenticity No jailbreak payloads. No special tool...

5.7AI score
SaveExploits0References3
Malwarebytes
Malwarebytes
•added 2026/09/21 2:21 p.m.•8 views

Gemini’s breach of real companies exposes an AI guardrail problem

Google says one of its Gemini models accessed systems belonging to three real companies during a cybersecurity evaluation in May. The model reportedly guessed credentials in one case, while finding exposed credentials in public repositories in two others. Google says Gemini stopped once it...

5.8AI score
SaveExploits0
Rows per page
Query Builder