2 matches found
better_opts_attacks
¿Puedo tener su atención? Rompiendo defensas de inyección de prompts basadas en fine-tuning mediante ataques conscientes de la arquitectura Este repositorio contiene el código para ejecutar los ataques ASTRA y ASTRA++ que rompen SecAlign++, SecAlign, StruQ. Este repositorio también contiene algun...
6.2AI score
SaveExploits0
May I Have Your Attention? Breaking Fine-Tuning Based Prompt Injection Defenses Using Architecture-Aware Attacks
A popular class of defenses against prompt injection attacks on large language models LLMs relies on fine-tuning the model to separate instructions and data, so that the LLM does not follow instructions that might be present with data. There are several academic systems and production-level...
7.4AI score
SaveExploits0
20