arXiv AI By Hanyang Zhao, Haoxian Chen, Han Lin, Genta Indra Winata, David Yao, Wenpin Tang

OPD+: Rethinking the Advantage Design for On-Policy Distillation

Read the original on arXiv AI →

arXiv:2606. 01039v1 Announce Type: cross Abstract: On-policy distillation (OPD) is a widely used technique to transfer capabilities from capable teacher language models to the base student models, and can be formulated in a reinforcement learning style objective using student generated rollouts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives

The paper investigates on‑policy distillation (OPD) versus supervised fine‑tuning (SFT), focusing on how students learn from multiple teachers by minimizing divergence. It shows that using forward KL divergence leads to a weighted arithmetic mixture, while reverse KL produces a normalized weighted geometric aggregate. The authors develop algorithms for both off‑policy and on‑policy settings, prove logarithmic regret bounds in tabular cases, extend the analysis to function approximation, and analyze how these aggregation targets explain OPD’s benefits and fragility.

By Qiwei Di, Xuheng Li, Kaixuan Ji, Chenggong Zhang, Heyang Zhao, Quanquan Gu
arXiv Machine Learning
Jul 7

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.

By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan