arXiv Machine Learning
Jun 30

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

arXiv:2606. 30406v1 Announce Type: cross Abstract: Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities, yet integrating multiple capabilities into one model remains hard.

By Wenhan Ma, Jianyu Wei, Liang Zhao, Hailin Zhang, Bangjun Xiao, Lei Li, Qibin Yang, Bofei Gao, Yudong Wang, Rang Li, Jinhao Dong, Zhifang Sui, Fuli Luo
arXiv Machine Learning
3d ago

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Latent-MOPD is a new on‑policy distillation method that allows a single large language model student to learn from multiple specialist teachers by using both the teachers’ output distributions and their hidden state representations. The approach selects late‑layer targets based on teacher‑student relationships, bridges hidden width differences with a shared projection, and groups updates by domain, enabling gradual shift from representation to token supervision. Experiments show that Latent‑MOPD outperforms token‑only, representation‑only, and uniform‑averaging baselines across nine benchmarks in math, code, and logic, and even surpasses the best individual teacher on most tasks.

By Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li
arXiv AI
Jun 30

DOPD: Dual On-policy Distillation

arXiv:2606. 30626v1 Announce Type: new Abstract: On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals.

By Xinlei Yu, Gen Li, Qingyi Si, Guibin Zhang, Yuqi Xu, Congcong Wang, Shuai Dong, Kaiwen Tuo, Xiangyu Zeng, Kaituo Feng, Qunzhong Wang, Yang Shi, Xiaobin Hu, Xiangyu Yue, Jiaqi Wang, Shuicheng Yan
arXiv Computation and Language
Aug 28

Preserving General Capabilities during Domain Specialization with Uncertainty-Calibrated MOPD

The paper introduces uncertainty‑calibrated Multi‑Teacher On‑Policy Distillation (MOPD) to better preserve general language model capabilities while specializing to specific domains. By employing dual‑temperature sampling, positive‑advantage‑density filtering, and centered log‑likelihood filtering, the method selects more informative trajectories and token updates, leading to significant improvements in general‑capability performance on role‑playing and medical‑domain tasks without sacrificing domain performance.

By Ziyuan Liu, Jiao Ou, Jian Liang, Ruiming Tang, Cheng Luo
arXiv Machine Learning
Jul 7

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.

By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan