arXiv AI
Sep 28

TISD: On-Policy Self-Distillation with Trajectory Intervention

The paper introduces TISD, a trajectory-intervention self-distillation method that forces a teacher-selected branch action and then lets the student generate the suffix, distilling the full trajectory under a privileged-context-conditioned teacher. This approach addresses a data-collection bottleneck in on‑policy self‑distillation by exposing successor contexts that the student would otherwise miss. Experiments on coding and science domains show modest but consistent improvements in average performance metrics compared to baseline methods.

By Taeckyung Lee, Rinat Amankos, Jeonghye Kim, Hyungjun Yoon, Woogyeol Jin, Sung-Ju Lee
arXiv AI
Aug 18

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

arXiv:2608. 16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories.

By Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu