3 matches found
better_opts_attacks
May I have your attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks This repository contains code to run the ASTRA and ASTRA++ attacks that break SecAlign++, SecAlign, StruQ. This repository also contains some examples of generated attacks and attack...
Meta_SecAlign
Meta SecAlign: Un LLM fundacional seguro contra ataques de inyección de prompts Sizhe Chen, Arman Zharmagambetov, David Wagner, Chuan Guo por contribuciones técnicas equivalentes 🔥 Los modelos Meta-SecAlign ahora cuentan con licencia para uso comercial bajo las licencias de la comunidad Llama, a...
May I Have Your Attention? Breaking Fine-Tuning Based Prompt Injection Defenses Using Architecture-Aware Attacks
A popular class of defenses against prompt injection attacks on large language models LLMs relies on fine-tuning the model to separate instructions and data, so that the LLM does not follow instructions that might be present with data. There are several academic systems and production-level...