arXiv AI

Tail-Aware Top-$k$ On-Policy Distillation

arXiv:2608. 14728v1 Announce Type: cross Abstract: On-policy distillation (OPD) has emerged as an effective paradigm for transferring knowledge between language models, where a student is trained to align its next-token distribution with the teacher's along its own trajectories.

arXiv Machine Learning
Sep 1

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

The paper investigates on‑policy distillation (OPD), showing that teacher supervision during OPD contains significant noise that grows with teacher size, yet the student policy remains largely unaffected by this noise. It finds that OPD’s gains stem mainly from suppressing low‑log‑probability tokens, a process that can be replicated without a teacher. Building on this insight, the authors propose On‑Policy Self‑Adaptation (OPSA), a supervision‑free method that uses entropy‑adaptive negative advantages to improve performance on several benchmarks, outperforming both the base model and OPD.

By Yi Ding, Ruqi Zhang
arXiv AI
3d ago

Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives

The paper investigates on‑policy distillation (OPD) versus supervised fine‑tuning (SFT), focusing on how students learn from multiple teachers by minimizing divergence. It shows that using forward KL divergence leads to a weighted arithmetic mixture, while reverse KL produces a normalized weighted geometric aggregate. The authors develop algorithms for both off‑policy and on‑policy settings, prove logarithmic regret bounds in tabular cases, extend the analysis to function approximation, and analyze how these aggregation targets explain OPD’s benefits and fragility.

By Qiwei Di, Xuheng Li, Kaixuan Ji, Chenggong Zhang, Heyang Zhao, Quanquan Gu
arXiv Machine Learning
Aug 27

A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation

The paper investigates how the sampled-token reverse-KL loss in on‑policy distillation distributes updates across tokens. By analyzing the gradient of the per‑token K2 estimator, the authors find that tokens with low student probability and large teacher‑student gaps receive disproportionately large gradient norms. They propose Surprise‑aware Reweighting (SuRe), a lightweight weighting rule that further amplifies this allocation, and demonstrate that SuRe improves math metrics on Qwen3 student models without harming out‑of‑domain performance.

By Bing Shao, Jiazheng Zhang, Long Ma, Yujiong Shen, Senjie Jin, Xin Guo, Yuming Yang, Mingxu Chai, Zhiheng Xi, Tao Gui, Qi Zhang, Xuanjing Huang