arXiv:2605.06387v4 Announce Type: replace-cross
Abstract: On-policy distillation (OPD) trains a student on its own trajectories with token-level teacher feedback and often outperforms off-policy dist...
By Nan Jia, Haojin Yang, Xing Ma, Jiesong Lian, Shuailiang Zhang, Weipeng Zhang, Ke Zeng, Xunliang Cai, Zequn Sun
arXiv:2609.38025v1 Announce Type: cross
Abstract: On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OP...
By Zhenyu Wang, Tianze Wang, Linjun Zhang, Yifan Hu
The paper investigates on‑policy distillation (OPD) versus supervised fine‑tuning (SFT), focusing on how students learn from multiple teachers by minimizing divergence. It shows that using forward KL divergence leads to a weighted arithmetic mixture, while reverse KL produces a normalized weighted geometric aggregate. The authors develop algorithms for both off‑policy and on‑policy settings, prove logarithmic regret bounds in tabular cases, extend the analysis to function approximation, and analyze how these aggregation targets explain OPD’s benefits and fragility.
By Qiwei Di, Xuheng Li, Kaixuan Ji, Chenggong Zhang, Heyang Zhao, Quanquan Gu
arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.
By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan
On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and optimization stability...
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that...