Lucene search
+L

2 matches found

Kitploit
Kitploit
•added 2026/10/05 5:03 p.m.•14 views

SecOPD

SecOPD: 通过在线策略蒸馏缓解自适应提示注入 Yibo Peng · Long Lian · David Wagner† · Sizhe Chen† † 联合指导。 论文 项目主页 模型 本版本实现了论文中最终的 全响应 KL 公式,也即 无解析 变体。学生模型在受攻击的上下文下进行 rollout。一个由相同基础模型初始化的干净上下文教师模型,在配对的干净上下文下对学生模型采样的 token 进行评分。反向 KL 信号应用于每个采样的响应 token,包括推理 token;不使用 边界来选择监督跨度。 仓库结构 training/tinker/:全响应 KL...

6.2AI score
SaveExploits0References3
Packet Storm News
Packet Storm News
•added 2026/09/26 12:00 a.m.•10 views

Leaky Students: Membership Inference against On-Policy Distillation

On-policy distillation OPD trains a student to match a teacher's next-token distributions on student-generated trajectories. However, privileged information supplied to the teacher for OPD training may contain sensitive data. Whether the student leaks private information about the records supplie...

5.8AI score
SaveExploits0
Rows per page
Query Builder