Lucene search
+L

4 matches found

Kitploit
Kitploit
•added 2026/10/01 11:57 a.m.•11 views

PoisonCraft

PoisonCraft Este repositorio proporciona la implementación oficial de POISONCRAFT: Envenenamiento Práctico de la Generación Aumentada por Recuperación para Modelos de Lenguaje de Gran Tamaño. Información General POISONCRAFT tiene como objetivo demostrar cómo un actor malicioso puede plantar...

6.3AI score
SaveExploits0References5
Packet Storm News
Packet Storm News
•added 2026/02/06 12:00 a.m.•36 views

TrapSuffix: Proactive Defense against Adversarial Suffixes in Jailbreaking

Suffix-based jailbreak attacks append an adversarial suffix, i.e., a short token sequence, to steer aligned LLMs into unsafe outputs. Since suffixes are free-form text, they admit endlessly many surface forms, making jailbreak mitigation difficult. Most existing defenses depend on passive detecti...

5.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/05/21 12:00 a.m.•17 views

Alignment under Pressure: the Case for Informed Adversaries When Evaluating LLM Defenses

Large language models LLMs are rapidly deployed in real-world applications ranging from chatbots to agentic systems. Alignment is one of the main approaches used to defend against attacks such as prompt injection and jailbreaks. Recent defenses report near-zero Attack Success Rates ASR even again...

6.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/05/14 12:00 a.m.•18 views

Adversarial Suffix Filtering: a Defense Pipeline for LLMs

Large Language Models LLMs are increasingly embedded in autonomous systems and public-facing environments, yet they remain susceptible to jailbreak vulnerabilities that may undermine their security and trustworthiness. Adversarial suffixes are considered to be the current state-of-the-art...

7AI score
SaveExploits0
Rows per page
Query Builder