arXiv Machine Learning

REVO: Rollout-Efficient Off-Policy Distillation via Variance-Guided Reuse

arXiv AI
Jun 9

Trajectory-Refined Distillation

arXiv:2606. 08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts.

By Li Jiang, Haoran Xu, Yichuan Ding, Amy Zhang
arXiv AI
Aug 18

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

arXiv:2608. 16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories.

By Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu
arXiv Machine Learning
Jul 7

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.

By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan
arXiv Machine Learning
Sep 22

Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation

Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation explores how on‑policy distillation (OPD) can be improved by addressing teacher uncertainty that contracts as the teacher continues from a student‑generated prefix. The authors identify Teacher Uncertainty Contraction (TUC) and theoretically analyze its variance‑bias trade‑off, leading to the proposal of Adaptive‑Continuations On‑Policy Distillation (AC‑OPD). Experiments on mathematical reasoning and code generation show that AC‑OPD consistently outperforms standard OPD, with controlled‑continuation and matched‑budget analyses supporting the adaptive‑continuation design.

By Jingang Zhou, Yuyi Zhou, Haiyang Guo, Xukai Wang, Shuai Feng, Sirui Gao, Jian Xu, Qingpei Guo, Xu-Yao Zhang
arXiv AI
3d ago

TISD: On-Policy Self-Distillation with Trajectory Intervention

The paper introduces TISD, a trajectory-intervention self-distillation method that forces a teacher-selected branch action and then lets the student generate the suffix, distilling the full trajectory under a privileged-context-conditioned teacher. This approach addresses a data-collection bottleneck in on‑policy self‑distillation by exposing successor contexts that the student would otherwise miss. Experiments on coding and science domains show modest but consistent improvements in average performance metrics compared to baseline methods.

By Taeckyung Lee, Rinat Amankos, Jeonghye Kim, Hyungjun Yoon, Woogyeol Jin, Sung-Ju Lee
arXiv AI
Sep 17

Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers

The paper introduces a method for offline on‑policy distillation that addresses the problem of imperfect teacher supervision. By training on teacher‑successful problems and measuring changes in token likelihoods on teacher‑failed trajectories, the authors derive a learnability signal that weights the distillation loss. This approach improves performance on mathematical reasoning and code generation tasks while reducing computational cost compared to online distillation.

By Yihao Ai, Weilong Yan