Lucene search
+L

1764 matches found

Kitploit
Kitploit
•added 2026/10/10 12:59 p.m.•18 views

inspect_petri

Inspect Petri Welcome to Inspect Petri, an auditing agent that enables automated monitoring and interaction with language models to detect potential alignment issues, reward hacking, and other concerning behaviors. Petri helps you rapidly test concrete alignment hypotheses end‑to‑end. It: Generat...

6.2AI score
SaveExploits0References2
Kitploit
Kitploit
•added 2026/10/10 9:48 a.m.•22 views

heretic

Heretic: Fully automatic censorship removal for language models Heretic is a tool that removes censorship aka "safety alignment" from transformer-based language models without expensive post-training. It combines an advanced implementation of directional ablation, also known as "abliteration"...

6AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/10 4:43 a.m.•17 views

permanently-jailbroken

Permanently Jailbroken We asked GPT-4, Claude, Gemini, DeepSeek, Grok, and Mistral 5 questions about their own programming. All 6 said jailbreaking will never be fixed. Not because the patches are bad. Because alignment doesn't change what the model understands — it changes what the model says. T...

6.3AI score
SaveExploits0References6
Kitploit
Kitploit
•added 2026/10/09 1:35 p.m.•12 views

RLCDAlignBench

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures This repository holds RLCDAlignBench and the code behind the paper. The benchmark measures whether a detector can tell when a language model's output is an alignment failure. It has 44...

6.3AI score
SaveExploits0References13
Kitploit
Kitploit
•added 2026/10/09 8:31 a.m.•20 views

redteam-plan

🔥 🚒 레드 팀 모의훈련 계획 이 문서는 Red Teams에 설명된 매우 구체적인 레드 팀 스타일과 대비하여 레드 팀 계획을 알리는 데 도움을 줍니다. 이 방법은 블루 팀의 가치와 열정을 최적화하기 위해 몇 가지 편향을 표현합니다. 특히 레드 팀의 처벌을 통해 동기를 부여하려는 시도를 피합니다. 아래 질문들을 검토하여 레드 팀 계획이 블루 팀의 가치를 위해 충분히 고려되었는지 테스트해 보세요. ❌ 부정적 동기 다음은 레드 팀 모의훈련을 추진하는 일반적인 이유입니다. 이는 사기나 팀 결속에 해로운 영향을 미칩니다. 모의훈련이 목...

6.2AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/09 12:42 a.m.•11 views

dod-blue-team-network-lab

DoD Blue Team Network Security & Hardening Lab Executive Summary High fidelity defensive security lab simulating a DoD aligned enterprise network. This project demonstrates structured network segmentation, system hardening, centralized telemetry ingestion, detection engineering, adversary...

6.3AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/10/06 4:55 a.m.•17 views

Rubrics-as-an-Attack-Surface

攻撃対象としてのルーブリック: LLM審査員におけるステルスな選好ドリフト 📊 データセット • 🤖 学習済みモデル • 📝 論文 • 💻 リポジトリ このリポジトリには、Ruomeng Ding、Yifei Pang、He Sun、Yizhong Wang、Steven Wu、Zhun Dengによる論文攻撃対象としてのルーブリック: LLM審査員におけるステルスな選好ドリフトのコードが含まれています。...

6.3AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/05 5:04 p.m.•9 views

fish-live-in-trees

Fish Live in Trees 本番LLMランタイム・アラインメント・コンテキスト・インジェクション RACI 著者: Sean Kavanagh 日付: 2026-02-16 環境: Gemini 3 Flash 無料枠 - 公開本番インターフェース エクスプロイト種別: コンテキスト・ピボット 事実的正確性 → 対人的真正性 ジェイルブレイク・ペイロードなし。 特別なツールなし。 ただリフレーミングするだけ。 AIを壊す最善の方法は、すでに壊れているとAIを説得することだ! 中核的発見...

5.8AI score
SaveExploits0References3
AstraLinux
AstraLinux
•added 2026/10/05 11:39 a.m.•16 views

Astra Linux – Vulnerability in Linux 6.1

In the Linux kernel, the following vulnerabilities have been resolved: arm64: Do not call NULL in docompatalignmentfixup. doalignmentt32tohandler only fixes alignment faults for specific instructions; otherwise, it returns NULL e.g., for LDREX. When this occurs, a signal is sent to the caller...

5.5CVSS6.2AI score0.00204EPSS
SaveExploits0References2
AstraLinux
AstraLinux
•added 2026/10/05 11:39 a.m.•10 views

Astra Linux – Vulnerability in Linux 6.1

In the Linux kernel, the following vulnerabilities have been resolved: device-dax: The pgoff alignment in daxsetmapping should use ALIGNDOWN instead of ALIGN. Otherwise, vmf-address, which is not aligned with faultsize, will be aligned to the next alignment, which can lead to memory failures due ...

5.5CVSS6.6AI score0.00265EPSS
SaveExploits0References2
AstraLinux
AstraLinux
•added 2026/10/05 11:39 a.m.•11 views

Astra Linux – Vulnerability in Linux 5.15

In the Linux kernel, the following vulnerability has been resolved: media: imx-jpeg: Buffer size aligned upwards. The hardware can support any image size WxH, with arbitrary W image width and H image height dimensions. The buffer size is aligned upwards for both the encoder and the decoder. The...

7.8CVSS6.3AI score0.00252EPSS
SaveExploits0References1
AstraLinux
AstraLinux
•added 2026/10/05 11:39 a.m.•12 views

Astra Linux – Vulnerability found in Linux 5.10, Linux 6.1, and Linux 5.15

In the Linux kernel, the following vulnerabilities have been resolved: i40e: added validation for the ringlen parameter. The ringlen parameter provided by the virtual function VF is assigned directly to the hardware memory context HMC without any validation. To address this issue, a upper boundar...

8.8CVSS6.2AI score0.00154EPSS
SaveExploits0References2
Kitploit
Kitploit
•added 2026/10/05 5:39 a.m.•11 views

gemini-2.5-pro-nf-tables-red-teamin

gemini-2.5-pro-nf-tables-red-teamin Google Gemini 2.5 Pro의 안전 정렬 정책, 가드레일, 그리고 레거시 Linux 커널 취약점 프리미티브CVE-2023-32233와 관련된 거부 동작 변화를 기록한 기술 사례 연구 및 타임라인 데이터셋입니다. 사례 연구 웹사이트 링크 : https://destawell.github.io/gemini-2.5-pro-nf-tables-red-teamin/...

7.8CVSS6.9AI score0.12966EPSS
SaveExploits8References1
Packet Storm News
Packet Storm News
•added 2026/10/01 12:00 a.m.•13 views

High-Quality Data Do Not Mean Safe! Poisoning LLMs after Data Selection

Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples. However, prior studies typically assume that poisoned samples directly enter downstream fine-tuning, overlooking quality-based selection in practical training pipelines. To fi...

5.9AI score
SaveExploits0
RedhatCVE
RedhatCVE
•added 2026/09/30 8:53 a.m.•14 views

CVE-2026-97989

A flaw was found in the Linux kernel's vDPA Device in Userspace VDUSE subsystem. The driver does not properly validate virtqueue alignment parameters during device configuration, allowing invalid or zero alignment values to pass into internal queue creation routines. A local user could exploit th...

5.5CVSS5.5AI score0.00155EPSS
SaveExploits0References4
Packet Storm News
Packet Storm News
•added 2026/09/29 12:00 a.m.•9 views

SURE: Framework for Safety to Construct Trustworthy AI

Warning: This paper contains harmful and offensive text. Recently, large language models such as GPT-4, and Claude have revolutionized tasks in various domains. As the use of these large language models increases, people are increasingly concerned about AI safety and demand that large language...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/09/29 12:00 a.m.•12 views

CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition

Text-to-image T2I models have substantially improved in language understanding, in-image text rendering, and visual composition, while their safety mechanisms do not always keep pace with these capabilities. This creates a cross-modal attack surface in which harmful semantics can remain...

5.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/09/28 12:00 a.m.•9 views

From Sharp Eyes to Expert Mind: Internalizing Expert Knowledge in MLLMs for Tampered Text Detection

Tampered Text Detection TTD is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective at capturing subtle manipulation traces but often generalize poorly across diverse document domains, while Multimodal Large Language Models MLLMs...

5.7AI score
SaveExploits0
CVE
CVE
•added 2026/09/25 10:23 a.m.•27 views

CVE-2026-97989

The Linux kernel contains a vulnerability in vduse where vduse_validate_config() fails to properly validate the vq_align parameter, checking only the upper bound. This allows invalid values to reach vring_create_virtqueue_map(). Specifically, because split-ring helpers use align - 1 as a bit mask...

6.1AI score0.00155EPSS
SaveExploits0References2
OSV
OSV
•added 2026/09/24 5:17 p.m.•7 views

AZL-103628 CVE-2026-97417 affecting package kernel 6.6.157.1-1

In the Linux kernel, the following vulnerability has been resolved: netfilter: nfconntrack: use getunalignedbe32 in tcpsack The timestamp-only fast path dereferences the option stream as be32 ptr, which assumes 4-byte alignment that the TCP option stream does not guarantee. Use getunalignedbe32...

7.5CVSS5.8AI score0.0049EPSS
SaveExploits0References1
Rows per page
Query Builder