Lucene search
+L

4 matches found

Kitploit
Kitploit
added 2026/09/09 11:15 p.m.4 views

PoisonCraft

PoisonCraft This repository provides the official implementation of POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models. Overview POISONCRAFT aims to demonstrate how a malicious actor can plant “poisoned” content into the corpus used by Retrieval-Augmented...

6AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/02/06 12:00 a.m.34 views

TrapSuffix: Proactive Defense against Adversarial Suffixes in Jailbreaking

Suffix-based jailbreak attacks append an adversarial suffix, i.e., a short token sequence, to steer aligned LLMs into unsafe outputs. Since suffixes are free-form text, they admit endlessly many surface forms, making jailbreak mitigation difficult. Most existing defenses depend on passive detecti...

5.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/21 12:00 a.m.13 views

Alignment under Pressure: the Case for Informed Adversaries When Evaluating LLM Defenses

Large language models LLMs are rapidly deployed in real-world applications ranging from chatbots to agentic systems. Alignment is one of the main approaches used to defend against attacks such as prompt injection and jailbreaks. Recent defenses report near-zero Attack Success Rates ASR even again...

6.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/14 12:00 a.m.14 views

Adversarial Suffix Filtering: a Defense Pipeline for LLMs

Large Language Models LLMs are increasingly embedded in autonomous systems and public-facing environments, yet they remain susceptible to jailbreak vulnerabilities that may undermine their security and trustworthiness. Adversarial suffixes are considered to be the current state-of-the-art...

7AI score
SaveExploits0
Rows per page
Query Builder