2 matches found
better_opts_attacks
May I have your attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks This repository contains code to run the ASTRA and ASTRA++ attacks that break SecAlign++, SecAlign, StruQ. This repository also contains some examples of generated attacks and attack...
6AI score
SaveExploits0
May I Have Your Attention? Breaking Fine-Tuning Based Prompt Injection Defenses Using Architecture-Aware Attacks
A popular class of defenses against prompt injection attacks on large language models LLMs relies on fine-tuning the model to separate instructions and data, so that the LLM does not follow instructions that might be present with data. There are several academic systems and production-level...
7.4AI score
SaveExploits0
20