Lucene search
+L

31 matches found

Kitploit
Kitploit
added 2026/09/23 2:02 a.m.13 views

inspect_petri

Inspect Petri Welcome to Inspect Petri, an auditing agent that enables automated monitoring and interaction with language models to detect potential alignment issues, reward hacking, and other concerning behaviors. Petri helps you rapidly test concrete alignment hypotheses end‑to‑end. It: Generat...

6AI score
SaveExploits0References2
Kitploit
Kitploit
added 2026/09/23 12:25 a.m.9 views

socbench

socbench Benchmark frontier reasoning LLMs as SOC agents on raw NetFlow data. socbench benchmarks frontier reasoning models as SOC agents: each model runs a bounded multi-turn agent loop against a deterministic, pre-indexed NetFlow corpus, with persona-scoped read-only tools, fixed dollar caps pe...

6.1AI score
SaveExploits0References1
Kitploit
Kitploit
added 2026/09/22 4:55 a.m.4 views

INTACT

SoK 다중 턴 탈옥 실험 이 저장소는 SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks의 메커니즘 분석에 사용된 코드의 릴리스 버전을 포함합니다. 저장소는 논문 부록 A에 설명된 실험에 필요한 코드, 프롬프트, 구성 파일 및 데이터셋만 유지하도록 정리되었습니다. 생성된 로그, 캐시된 파일, 가상 환경, 이전 결과, 그림 및 임시 분석 출력은 의도적으로 제거되었습니다. 실험 매핑 논문 질문| 방법| 범주| 로컬 디렉터리| 데이터셋...

5.8AI score
SaveExploits0
Kitploit
Kitploit
added 2026/09/20 2:33 a.m.6 views

PE-CoA

PE-CoA Код реализации статьи «Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models» Полная версия статьи доступна по ссылке: https://arxiv.org/pdf/2510.08859 Инструкции по настройке Chain of Attack Установка 1. Установите зависимости :...

6.1AI score
SaveExploits0
GithubExploit
GithubExploit
added 2026/07/26 11:21 p.m.81 views

DarkPrompt

DarkPrompt !CIhttps://github.com/jason-allen-oneal/DarkPr...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/30 12:00 a.m.44 views

Quality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM Safety

Current approaches to LLM adversarial testing suffer from coverage gaps: manual red-teaming does not scale, LLM-as-attacker methods exhibit mode collapse, and gradient-based approaches produce uninterpretable gibberish. We introduce a quality-diversity evolutionary framework that operates at the...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/09 12:00 a.m.37 views

MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks

Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can gradually steer the conversation toward an unsafe answer. Recent methods demonstrate this risk, but they are usually...

5.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/04/27 12:00 a.m.14 views

Jailbreaking Frontier Foundation Models through Intention Deception

Large vision-language models exhibit remarkable capability but remain highly susceptible to jailbreaking. Existing safety training approaches aim to have the model learn a refusal boundary between safe and unsafe, based on the user's intent. It has been found that this binary training regime ofte...

5.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/04/07 12:00 a.m.14 views

Your LLM Agent Can Leak Your Data: Data Exfiltration Via Backdoored Tool Use

Tool-use large language model LLM agents are increasingly deployed to support sensitive workflows, relying on tool calls for retrieval, external API access, and session memory management. While prior research has examined various threats, the risk of systematic data exfiltration by backdoored...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/01/09 12:00 a.m.17 views

The Echo Chamber Multi-Turn LLM Jailbreak

The availability of Large Language Models LLMs has led to a new generation of powerful chatbots that can be developed at relatively low cost. As companies deploy these tools, security challenges need to be addressed to prevent financial loss and reputational damage. A key security challenge is...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/01/08 12:00 a.m.45 views

Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models

Large Language Models LLMs face a significant threat from multi-turn jailbreak attacks, where adversaries progressively steer conversations to elicit harmful outputs. However, the practical effectiveness of existing attacks is undermined by several critical limitations: they struggle to maintain ...

7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/01/08 12:00 a.m.11 views

Multi-Turn Jailbreaking Attack in Multi-Modal Large Language Models

In recent years, the security vulnerabilities of Multi-modal Large Language Models MLLMs have become a serious concern in the Generative Artificial Intelligence GenAI research. These highly intelligent models, capable of performing multi-modal tasks with high accuracy, are also severely susceptib...

7.2AI score
SaveExploits0
HackRead
HackRead
added 2025/11/11 10:35 a.m.15 views

Cisco Finds Open-Weight AI Models Easy to Exploit in Long Chats

Cisco’s new research shows that open-weight AI models, while driving innovation, face serious security risks as multi-turn attacks, including conversational persistence, can bypass safeguards and expose data...

7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/11/04 12:00 a.m.18 views

AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models

Large Language Models LLMs remain vulnerable to jailbreaking attacks where adversarial prompts elicit harmful outputs, yet most evaluations focus on single-turn interactions while real-world attacks unfold through adaptive multi-turn conversations. We present AutoAdv, a training-free framework fo...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/10/21 12:00 a.m.20 views

HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models

Large Language Models LLMs remain vulnerable to multi-turn jailbreak attacks. We introduce HarmNet, a modular framework comprising ThoughtNet, a hierarchical semantic network; a feedback-driven Simulator for iterative query refinement; and a Network Traverser for real-time adaptive attack...

7.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/10/16 12:00 a.m.21 views

Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks

Large language models LLMs are increasingly vulnerable to multi-turn jailbreak attacks, where adversaries iteratively elicit harmful behaviors that bypass single-turn safety filters. Existing defenses predominantly rely on passive rejection, which either fails against adaptive attackers or overly...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/10/09 12:00 a.m.15 views

Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models

Large language models LLMs remain vulnerable to multi-turn jailbreaking attacks that exploit conversational context to bypass safety constraints gradually. These attacks target different harm categories like malware generation, harassment, or fraud through distinct conversational approaches...

7.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/10/03 12:00 a.m.11 views

NEXUS: Network Exploration for EXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks

Large Language Models LLMs have revolutionized natural language processing but remain vulnerable to jailbreak attacks, especially multi-turn jailbreaks that distribute malicious intent across benign exchanges and bypass alignment mechanisms. Existing approaches often explore the adversarial space...

7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/09/29 12:00 a.m.15 views

STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents

As LLMs advance into autonomous agents with tool-use capabilities, they introduce security challenges that extend beyond traditional content-based LLM safety concerns. This paper introduces Sequential Tool Attack Chaining STAC, a novel multi-turn attack framework that exploits agent tool use. STA...

7.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/09/04 12:00 a.m.13 views

Between a Rock and a Hard Place: Exploiting Ethical Reasoning to Jailbreak LLMs

Large language models LLMs have undergone safety alignment efforts to mitigate harmful outputs. However, as LLMs become more sophisticated in reasoning, their intelligence may introduce new security risks. While traditional jailbreak attacks relied on singlestep attacks, multi-turn jailbreak...

7.4AI score
SaveExploits0
Rows per page
Query Builder