1 matches found
Backdoor-Attack-Defense-LLMs
Interpretability-of-LLMs IBSDは、IEEE Signal Processing Letters(2025.10)に掲載された論文「IBSD: Iterable Black-box Self-defense Against Backdoor Attacks」のコードです。論文リンク SLIPは、2026-ACL-findingsに掲載された論文「SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in...
5.8AI score
SaveExploits0
20