arXiv AI By Tianze Xu, Yanzhao Zheng, Zhentao Zhang, Yuanqiang Yu, Chao Ma, Jihuai Zhu, Lelun Wu, Lyumanshan Ye, Pengfei Liu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation

Read the original on arXiv AI →

MOPD‑Router rethinks teacher routing in multi‑teacher on‑policy distillation by routing supervision over the full teacher pool at each token, eliminating the need for prompt‑level domain labels or a separate routing model. The framework offers a plug‑in interface for various metrics, and introduces ExpertAlign, which scores teachers based on how well their corrections reflect their specialized post‑training knowledge. Experiments on both unlabeled and domain‑labeled mixtures show that ExpertAlign outperforms existing methods, improving overall scores by up to 12.3% on unlabeled data and 7.8% on domain‑labeled data. whyItMatters":"Token‑level routing enables the use of complementary supervision across domains without relying on domain labels, leading to significant performance gains in multi‑teacher distillation settings."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 22

Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation

arXiv:2609.23697v1 Announce Type: cross Abstract: Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, howe...

By Jie Sun, Mao Zheng, Mingyang Song, Zeyuan Liu, Gengsheng Li, Houcheng Jiang, Yilin Cheng, Bichuan Feng, Yuchen Cai, Junfeng Fang, Xiang Wang
arXiv AI
Aug 18

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

arXiv:2608. 16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories.

By Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu
arXiv Machine Learning
Sep 1

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

The paper investigates on‑policy distillation (OPD), showing that teacher supervision during OPD contains significant noise that grows with teacher size, yet the student policy remains largely unaffected by this noise. It finds that OPD’s gains stem mainly from suppressing low‑log‑probability tokens, a process that can be replicated without a teacher. Building on this insight, the authors propose On‑Policy Self‑Adaptation (OPSA), a supervision‑free method that uses entropy‑adaptive negative advantages to improve performance on several benchmarks, outperforming both the base model and OPD.

By Yi Ding, Ruqi Zhang
arXiv Computation and Language
Aug 28

Preserving General Capabilities during Domain Specialization with Uncertainty-Calibrated MOPD

The paper introduces uncertainty‑calibrated Multi‑Teacher On‑Policy Distillation (MOPD) to better preserve general language model capabilities while specializing to specific domains. By employing dual‑temperature sampling, positive‑advantage‑density filtering, and centered log‑likelihood filtering, the method selects more informative trajectories and token updates, leading to significant improvements in general‑capability performance on role‑playing and medical‑domain tasks without sacrificing domain performance.

By Ziyuan Liu, Jiao Ou, Jian Liang, Ruiming Tang, Cheng Luo