arXiv:2609.33455v2 Announce Type: replace
Abstract: On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since e...
By Zizhuo Lin, Quanling Liu, Yi Yang, Yawei Luo
arXiv:2609.14636v1 Announce Type: new
Abstract: On-policy distillation (OPD) has become a standard approach for transferring capabilities from large teachers to compact students. Its cost, however, i...
By Zhiyu Gui, Kexin Huang, Jia Guo, Junkang Wu, Zihao Wang, Zhiqiang Zhang, Jun Zhou, Jiancan Wu, Xiang Wang
arXiv:2609.39687v1 Announce Type: new
Abstract: On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises studen...
By Xincheng Wei, Yifan Ding, Yoshua Li, Yuquan Lu, Ziheng Li, Yi Lu, Dongsheng Ma, Rongxiang Weng, Xunliang Cai
arXiv:2608. 08726v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) uses a privileged teacher to supervise a reasoning model on prefixes sampled from its own rollouts.
By Yangyang Feng, Zhuoyan Feng, Junlan Chen
arXiv:2606. 21994v2 Announce Type: replace Abstract: On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories.
By Qingfei Zhao, Huan Song, Shuyu Tian, Jiawei Shao, Xuelong Li
arXiv:2606. 08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts.
By Li Jiang, Haoran Xu, Yichuan Ding, Amy Zhang