arXiv Machine Learning By Qingfei Zhao, Huan Song, Shuyu Tian, Jiawei Shao, Xuelong Li

Prefix-Guided On-Policy Distillation: Mining Golden Trajectories from Rollouts

Read the original on arXiv Machine Learning →

arXiv:2606. 21994v2 Announce Type: replace Abstract: On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
3d ago

On the Off-Policy Teacher in On-Policy Distillation

The paper introduces SCOUT, a co‑training framework that adapts an off‑policy teacher to better continue from student‑generated prefixes in on‑policy distillation (OPD). By periodically optimizing the teacher’s conditional continuation ability using reinforcement learning with verifiable rewards, SCOUT improves the teacher’s performance on student prefixes. Experiments across various teacher‑student setups, model scales, and reasoning domains show that SCOUT consistently enhances the effectiveness of OPD.

By Langlin Huang, Hao Liu, Mononito Goswami, Xinyu Li, Prithwith Jana, Nikos Kanakaris, Patrick Bl\"obaum, Purak Jain
arXiv AI
Aug 18

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

arXiv:2608. 16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories.

By Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu