2 matches found
Prompt-Injection-in-the-Wild
Prompt Injection в реальных условиях Трекер публично задокументированных техник prompt injection примерно за последние два года, поддерживаемый Rachel James cybershujin. Каждая техника разбита на три элемента prompt injection: 1. Способ доставки — как внедрённые инструкции достигают модели...
6AI score
SaveExploits0References5
Mitigating Jailbreaks with Intent-Aware LLMs
Despite extensive safety-tuning, large language models LLMs remain vulnerable to jailbreak attacks via adversarially crafted instructions, reflecting a persistent trade-off between safety and task performance. In this work, we propose Intent-FT, a simple and lightweight fine-tuning approach that...
7.2AI score
SaveExploits0
20