arXiv Machine Learning By Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Read the original on arXiv Machine Learning →

Latent-MOPD is a new on‑policy distillation method that allows a single large language model student to learn from multiple specialist teachers by using both the teachers’ output distributions and their hidden state representations. The approach selects late‑layer targets based on teacher‑student relationships, bridges hidden width differences with a shared projection, and groups updates by domain, enabling gradual shift from representation to token supervision. Experiments show that Latent‑MOPD outperforms token‑only, representation‑only, and uniform‑averaging baselines across nine benchmarks in math, code, and logic, and even surpasses the best individual teacher on most tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 30

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

arXiv:2606. 30406v1 Announce Type: cross Abstract: Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities, yet integrating multiple capabilities into one model remains hard.

By Wenhan Ma, Jianyu Wei, Liang Zhao, Hailin Zhang, Bangjun Xiao, Lei Li, Qibin Yang, Bofei Gao, Yudong Wang, Rang Li, Jinhao Dong, Zhifang Sui, Fuli Luo
arXiv Machine Learning
Jun 5

OPRD: On-Policy Representation Distillation

arXiv:2606. 06021v1 Announce Type: new Abstract: On-policy distillation (OPD) supervises the student only in output space by matching next-token probabilities.

By Shenzhi Yang, Guangcheng Zhu, Bowen Song, Haobo Wang, Mingxuan Xia, Xing Zheng, Yingfan Ma, Zhongqi Chen, Weiqiang Wang, Gang Chen
arXiv AI
Sep 25

Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD

The paper introduces Selective Supervision for Direct-OPD (S$^2$D-OPD), a refinement of Direct On-Policy Distillation that filters out states where the teacher’s policy change is minimal, as measured by the teacher‑reference Jensen‑Shannon divergence. By masking low‑divergence states and keeping only the top 10% of states per response, S$^2$D-OPD improves held‑out accuracy on AIME and HMMT benchmarks across multiple teacher‑student pairs without additional forward passes.

By Yibo Zhao, Zixuan Yang, Yunshi Lan, Xiang Li