Lucene search
+L

545 matches found

Kitploit
Kitploit
•added 2026/10/04 2:40 p.m.•14 views

claude-XML-injection

Claude Sonnet 4.6 — Credential Fabrication via XML Tag Injection Reported: June 14, 2026 | Patched: June 20, 2026 | Anthropic Response: None 56 days Researcher: X1NON Affected Model: Claude Sonnet 4.6 and other non-Haiku Claude models Severity: High CVSS 8.7 Status: Patched — No bounty awarded, n...

6.2AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/10/04 2:52 a.m.•16 views

SparstanBoogie

SparstanBoogie se probó, según se informa, en iOS/iPadOS 15.2 - 16.7 RC 20H18 y 17.0; sin embargo, las pruebas se han realizado hasta 26.1 26.1. También se agradece cualquier prueba en dispositivos por encima o por debajo de lo informado, ya que podría volver a aparecer en dispositivos más nuevos...

6AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/10/04 2:33 a.m.•14 views

PE-CoA

PE-CoA Implementación en código de "Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models" El artículo completo está disponible en: https://arxiv.org/pdf/2510.08859 Instrucciones de configuración de Chain of Attack Instalación 1. Instalar...

6.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/10/04 12:00 a.m.•5 views

Which Image Property Carries the Jailbreak? A Controlled Dissection of Image-To-Text Jailbreaks

Image-to-text jailbreaks place harmful intent in text, image content, or the relationship between them. We examine image-side factors across four published attack families on a 313-prompt StrongREJECT slice, using five multimodal models and an additional appendix evaluation of InternVL3.5-8B. The...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/10/04 12:00 a.m.•6 views

Where Does the Audio Jailbreak Live? A Controlled Frequency-Depth Audit of AdvWave-P on Qwen2-Audio

We audit frequency and decoder-depth claims for AdvWave-P, an additive audio jailbreak, on Qwen2-Audio. The protocol masks frequency components of the perturbation in the short-time Fourier transform STFT domain and measures attack success and audio-span representations. On 520 AdvBench prompts,...

5.8AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/10/03 11:30 p.m.•14 views

skybreak

skybreak 8.4.1 Jailbreak usando CVE-2016-4655 / CVE-2016-4656 Crédito: Bellis1000 Billy Ellis, jndok...

9.3CVSS7.2AI score0.33353EPSS
SaveExploits8
Packet Storm News
Packet Storm News
•added 2026/10/03 12:00 a.m.•7 views

Reactivating Alignment: Defending LLMs from Jailbreaks Via Intention-Aware Input-Output Matching

Large language models LLMs remain vulnerable to jailbreak attacks that conceal harmful intent within complex adversarial prompts. Existing defenses primarily rely on input perturbation or harmful-output suppression, but they rarely model where malicious intent resides, resulting in brittle...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/10/01 12:00 a.m.•6 views

Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks

Recent work has shown that large language models LLMs can be vulnerable to jailbreak attacks in which harmful intent is obscured through composition with benign tasks. A harmful request refused in isolation may elicit a different response when embedded within a larger, seemingly benign query. We...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/09/30 12:00 a.m.•16 views

CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models Via Structured Code Completion

Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained predominantly on natural...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/09/29 12:00 a.m.•12 views

CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition

Text-to-image T2I models have substantially improved in language understanding, in-image text rendering, and visual composition, while their safety mechanisms do not always keep pace with these capabilities. This creates a cross-modal attack surface in which harmful semantics can remain...

5.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/09/29 12:00 a.m.•11 views

Does the Unsafe Gradient Survive a Conversation? on the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue

Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/09/29 12:00 a.m.•8 views

SceneJail: Exploiting Video Scenario Context to Jailbreak Multimodal LLMs

Video Multimodal Large Language Models Video-MLLMs support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby treating video merely as a...

5.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/09/29 12:00 a.m.•10 views

Controlled Decoding Attacks on Black-Box LLMs

Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/09/25 12:00 a.m.•15 views

TempQ-Jail: Query-Constrained Candidate Ranking for Text-To-Video Jailbreak Attacks

Existing text-to-video T2V jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video generation and security evaluation are costly, so an attacker often cannot test a large candidate pool. We therefore formulate T2V jailbreak as a...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/09/21 12:00 a.m.•7 views

Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection

Large language models LLMs are increasingly deployed in production systems, raising concerns about their exposure to adversarial manipulation through prompt injection and jailbreak attacks. Classifier-based guardrails, such as Prompt Guard 2, are widely used as a first line of defense against suc...

5.8AI score
SaveExploits0
Malwarebytes
Malwarebytes
•added 2026/09/18 2:18 p.m.•13 views

Did an AI really try to break free from human control?

Amid discussions about slowing down AI development, the Telegraph ran the headline: “OpenAI sounds alarm after bot tries to break free from human control.” That headline is slightly misleading, in my opinion. The Telegraph headline overstates what happened, although the underlying behavior is sti...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/09/18 12:00 a.m.•11 views

CASCADE against Jailbreaks: Combination across Stages with Controlled Attack-Defense Evaluation

Defenses against jailbreak attacks on Large Language Models LLMs operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine them. Prior empirical studies, fragmented by inconsistent...

5.8AI score
SaveExploits0
Kitploit
Kitploit
•added 2026/09/16 6:59 p.m.•14 views

augustus v0.14.30

Augustus - LLM vulnerability scanner for prompt injection, jailbreak, and adversarial attack testing Augustus - LLM Vulnerability Scanner Test large language models against 210+ adversarial attacks covering prompt injection, jailbreaks, encoding exploits, and data extraction. Augustus is a Go-bas...

6.1AI score
SaveExploits0References3
GithubExploit
GithubExploit
•added 2026/09/16 1:52 a.m.•31 views

Whetstone

Whetstone On-device JIT enabler for iOS 12 / 13, all devi...

7CVSS7.3AI score0.4608EPSS
SaveExploits10
Kitploit
Kitploit
•added 2026/09/15 10:45 p.m.•17 views

reasongate v0.4.0

ReasonGate A self-hostable gate that inspects the text going into and out of an LLM and returns an explainable allow / flag / block decision with a machine-readable audit record for every call. What this is The open-source core is rule-based. It does four things: recognizes known prompt-injection...

5.7AI score
SaveExploits0References2
Rows per page
Query Builder