arXiv AI

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation

MOPD‑Router rethinks teacher routing in multi‑teacher on‑policy distillation by routing supervision over the full teacher pool at each token, eliminating the need for prompt‑level domain labels or a separate routing model. The framework offers a plug‑in interface for various metrics, and introduces ExpertAlign, which scores teachers based on how well their corrections reflect their specialized post‑training knowledge. Experiments on both unlabeled and domain‑labeled mixtures show that ExpertAlign outperforms existing methods, improving overall scores by up to 12.3% on unlabeled data and 7.8% on domain‑labeled data. whyItMatters":"Token‑level routing enables the use of complementary supervision across domains without relying on domain labels, leading to significant performance gains in multi‑teacher distillation settings."

arXiv Machine Learning
Sep 22

Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation

arXiv:2609.23697v1 Announce Type: cross Abstract: Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, howe...

By Jie Sun, Mao Zheng, Mingyang Song, Zeyuan Liu, Gengsheng Li, Houcheng Jiang, Yilin Cheng, Bichuan Feng, Yuchen Cai, Junfeng Fang, Xiang Wang
arXiv AI
Aug 18

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

arXiv:2608. 16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories.

By Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu
arXiv Machine Learning
Sep 1

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

The paper investigates on‑policy distillation (OPD), showing that teacher supervision during OPD contains significant noise that grows with teacher size, yet the student policy remains largely unaffected by this noise. It finds that OPD’s gains stem mainly from suppressing low‑log‑probability tokens, a process that can be replicated without a teacher. Building on this insight, the authors propose On‑Policy Self‑Adaptation (OPSA), a supervision‑free method that uses entropy‑adaptive negative advantages to improve performance on several benchmarks, outperforming both the base model and OPD.

By Yi Ding, Ruqi Zhang
arXiv Computation and Language
Aug 28

Preserving General Capabilities during Domain Specialization with Uncertainty-Calibrated MOPD

The paper introduces uncertainty‑calibrated Multi‑Teacher On‑Policy Distillation (MOPD) to better preserve general language model capabilities while specializing to specific domains. By employing dual‑temperature sampling, positive‑advantage‑density filtering, and centered log‑likelihood filtering, the method selects more informative trajectories and token updates, leading to significant improvements in general‑capability performance on role‑playing and medical‑domain tasks without sacrificing domain performance.

By Ziyuan Liu, Jiao Ou, Jian Liang, Ruiming Tang, Cheng Luo
arXiv AI
Sep 3

Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs

The paper introduces Multi-Teacher Self-Distillation Policy Optimization (MT‑SDPO), an on‑policy distillation method that combines multiple frozen teachers into a single student model. MT‑SDPO uses self‑anchors, answer‑verified eligibility, and privileged distillation to select reliable teachers per sample rather than per domain. Experiments on five students from three model families show that MT‑SDPO improves the weakest domain of Qwen3‑8B by 14.79 points and reduces its domain gap by 74.7%, achieving a more balanced performance than matching a single teacher to each domain.

By Xixiang He, Xingming Li, Baiqi Wu, Qiyao Sun, Xuanyu Ji, Ao Cheng, Qingyong Hu
arXiv AI
Jun 9

Trajectory-Refined Distillation

arXiv:2606. 08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts.

By Li Jiang, Haoran Xu, Yichuan Ding, Amy Zhang
arXiv AI
Sep 18

D$^3$-MOPD: Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

The paper introduces D$^3$-MOPD, a dynamic domain scheduling method for multi-teacher on‑policy distillation. It adapts the domain mixture during training by monitoring each domain’s reverse‑KL trajectory, thereby allocating more compute to slower‑converging domains and less to those that plateau early. Experiments on a Qwen3.6‑35B‑A3B student show that D$^3$-MOPD closes 97% of the student‑to‑teacher performance gap, matches peak performance with roughly three times fewer rollout steps, and outperforms specialist teachers on most benchmarks.

By Zechen Sun, Zhiwei Zhang, Fei Zhao, Juntao Li, Mu Chuan, Huayu Deng, Guojian Zhan, Wenliang Chen, Yao Hu, Min Zhang
arXiv Computation and Language
Aug 25

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

arXiv:2608.16647v2 Announce Type: replace Abstract: On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalizati...

By Zhaoyi Li, Deyang Kong, Yuan Wei, Evan Yang, Ranran Shen, Mahardika Krisna Ihsani, Ming Yang, Wei Zhang, Chuan Hao, Jian Yang, Ran Tao, Bryan Dai, Shikun Zhang, Wei Ye, Ying Wei, Defu Lian
arXiv Machine Learning
Aug 27

D$^3$-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

The paper introduces D$^3$-MOPD, a zero‑overhead scheduler that dynamically adjusts domain sampling ratios during multi‑teacher on‑policy distillation by monitoring per‑domain reverse‑KL signals. Unlike static mixtures, D$^3$-MOPD reallocates compute toward slower‑converging domains, improving efficiency and performance. In experiments with a Qwen3.6‑35B‑A3B student distilled from four domain‑expert teachers, the method closes 97% of the student‑to‑teacher gap, matches peak performance with roughly three times fewer rollout steps, and outperforms specialist teachers on most benchmarks.

By Zechen Sun, Zhiwei Zhang, Fei Zhao, Juntao Li, Mu Chuan, Huayu Deng, Guojian Zhan, Wenliang Chen, Yao Hu, Min Zhang