3 matches found
SLDR
NeurIPS 2026 SLDR: الدفاع ضد الضبط الدقيق الخبيث عبر استعادة الطبقات الانتقائية والتوجيه الديناميكي 🔧 إعداد البيئة استنسخ المستودع، وأنشئ بيئة Conda وقم بتفعيلها، ثم ثبّت التبعيات المطلوبة: git clone https://github.com/Stardust457/SLDR.git cd SLDR conda create -n sldr python=3.12 -y conda activat...
SDD: Self-Degraded Defense against Malicious Fine-Tuning
Open-source Large Language Models LLMs often employ safety alignment methods to resist harmful instructions. However, recent research shows that maliciously fine-tuning these LLMs on harmful data can easily bypass these safeguards. To counter this, we theoretically uncover why malicious fine-tuni...
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
Fine-tuning-as-a-service, while commercially successful for Large Language Model LLM providers, exposes models to harmful fine-tuning attacks. As a widely explored defense paradigm against such attacks, unlearning attempts to remove malicious knowledge from LLMs, thereby essentially preventing th...