arXiv AI

Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives

The paper investigates on‑policy distillation (OPD) versus supervised fine‑tuning (SFT), focusing on how students learn from multiple teachers by minimizing divergence. It shows that using forward KL divergence leads to a weighted arithmetic mixture, while reverse KL produces a normalized weighted geometric aggregate. The authors develop algorithms for both off‑policy and on‑policy settings, prove logarithmic regret bounds in tabular cases, extend the analysis to function approximation, and analyze how these aggregation targets explain OPD’s benefits and fragility.

arXiv AI
Sep 17

Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers

The paper introduces a method for offline on‑policy distillation that addresses the problem of imperfect teacher supervision. By training on teacher‑successful problems and measuring changes in token likelihoods on teacher‑failed trajectories, the authors derive a learnability signal that weights the distillation loss. This approach improves performance on mathematical reasoning and code generation tasks while reducing computational cost compared to online distillation.

By Yihao Ai, Weilong Yan
arXiv Machine Learning
Aug 27

A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation

The paper investigates how the sampled-token reverse-KL loss in on‑policy distillation distributes updates across tokens. By analyzing the gradient of the per‑token K2 estimator, the authors find that tokens with low student probability and large teacher‑student gaps receive disproportionately large gradient norms. They propose Surprise‑aware Reweighting (SuRe), a lightweight weighting rule that further amplifies this allocation, and demonstrate that SuRe improves math metrics on Qwen3 student models without harming out‑of‑domain performance.

By Bing Shao, Jiazheng Zhang, Long Ma, Yujiong Shen, Senjie Jin, Xin Guo, Yuming Yang, Mingxu Chai, Zhiheng Xi, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv AI
3d ago

On the Off-Policy Teacher in On-Policy Distillation

The paper introduces SCOUT, a co‑training framework that adapts an off‑policy teacher to better continue from student‑generated prefixes in on‑policy distillation (OPD). By periodically optimizing the teacher’s conditional continuation ability using reinforcement learning with verifiable rewards, SCOUT improves the teacher’s performance on student prefixes. Experiments across various teacher‑student setups, model scales, and reasoning domains show that SCOUT consistently enhances the effectiveness of OPD.

By Langlin Huang, Hao Liu, Mononito Goswami, Xinyu Li, Prithwith Jana, Nikos Kanakaris, Patrick Bl\"obaum, Purak Jain
arXiv AI
Jun 2

OPD+: Rethinking the Advantage Design for On-Policy Distillation

arXiv:2606. 01039v1 Announce Type: cross Abstract: On-policy distillation (OPD) is a widely used technique to transfer capabilities from capable teacher language models to the base student models, and can be formulated in a reinforcement learning style objective using student generated rollouts.

By Hanyang Zhao, Haoxian Chen, Han Lin, Genta Indra Winata, David Yao, Wenpin Tang
arXiv Machine Learning
Jul 7

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.

By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan
arXiv AI
Jun 9

Trajectory-Refined Distillation

arXiv:2606. 08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts.

By Li Jiang, Haoran Xu, Yichuan Ding, Amy Zhang
arXiv AI
Aug 18

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

arXiv:2608. 16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories.

By Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu