arXiv:2609.33455v2 Announce Type: replace
Abstract: On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since e...
By Zizhuo Lin, Quanling Liu, Yi Yang, Yawei Luo
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conve...
On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable.
arXiv:2609.37500v1 Announce Type: new
Abstract: On-policy distillation (OPD) trains language models using dense token-level teacher supervision on student-generated trajectories. However, its relianc...
By Yuxiao Yang, Shangzhe Li, Tianrun Yu, Kaixiang Zhao, Taylor W. Killian, Weitong Zhang
arXiv:2608. 16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories.
By Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu
The paper introduces TISD, a trajectory-intervention self-distillation method that forces a teacher-selected branch action and then lets the student generate the suffix, distilling the full trajectory under a privileged-context-conditioned teacher. This approach addresses a data-collection bottleneck in on‑policy self‑distillation by exposing successor contexts that the student would otherwise miss. Experiments on coding and science domains show modest but consistent improvements in average performance metrics compared to baseline methods.
By Taeckyung Lee, Rinat Amankos, Jeonghye Kim, Hyungjun Yoon, Woogyeol Jin, Sung-Ju Lee