arXiv Computer Vision

UP-MOPD: Update Projection in Multi-Teacher On-Policy Distillation

arXiv AI
Sep 18

D$^3$-MOPD: Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

The paper introduces D$^3$-MOPD, a dynamic domain scheduling method for multi-teacher on‑policy distillation. It adapts the domain mixture during training by monitoring each domain’s reverse‑KL trajectory, thereby allocating more compute to slower‑converging domains and less to those that plateau early. Experiments on a Qwen3.6‑35B‑A3B student show that D$^3$-MOPD closes 97% of the student‑to‑teacher performance gap, matches peak performance with roughly three times fewer rollout steps, and outperforms specialist teachers on most benchmarks.

By Zechen Sun, Zhiwei Zhang, Fei Zhao, Juntao Li, Mu Chuan, Huayu Deng, Guojian Zhan, Wenliang Chen, Yao Hu, Min Zhang
arXiv Machine Learning
Aug 27

D$^3$-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

The paper introduces D$^3$-MOPD, a zero‑overhead scheduler that dynamically adjusts domain sampling ratios during multi‑teacher on‑policy distillation by monitoring per‑domain reverse‑KL signals. Unlike static mixtures, D$^3$-MOPD reallocates compute toward slower‑converging domains, improving efficiency and performance. In experiments with a Qwen3.6‑35B‑A3B student distilled from four domain‑expert teachers, the method closes 97% of the student‑to‑teacher gap, matches peak performance with roughly three times fewer rollout steps, and outperforms specialist teachers on most benchmarks.

By Zechen Sun, Zhiwei Zhang, Fei Zhao, Juntao Li, Mu Chuan, Huayu Deng, Guojian Zhan, Wenliang Chen, Yao Hu, Min Zhang
arXiv Machine Learning
5d ago

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

The paper investigates how teacher signals influence parameter updates in Multi‑Teacher On‑Policy Distillation (MOPD) by analyzing Qwen3‑1.7B and SmolLM3‑3B. It shows that loss averaging, Adam’s first‑moment bias, BF16 rounding, and the choice of averaging rule all shape the gradients and ultimately affect task performance. The study quantifies these effects, revealing, for example, that token‑averaging favors longer responses and that BF16 rounding masks most weight changes.

By Siqi Zhu, Suozhi Huang, Kaixuan Zhang, Yuheng Yang, Zhanyang Jin, Yihang Sun, Jiaxuan You
arXiv Machine Learning
3d ago

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Latent-MOPD is a new on‑policy distillation method that allows a single large language model student to learn from multiple specialist teachers by using both the teachers’ output distributions and their hidden state representations. The approach selects late‑layer targets based on teacher‑student relationships, bridges hidden width differences with a shared projection, and groups updates by domain, enabling gradual shift from representation to token supervision. Experiments show that Latent‑MOPD outperforms token‑only, representation‑only, and uniform‑averaging baselines across nine benchmarks in math, code, and logic, and even surpasses the best individual teacher on most tasks.

By Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li
arXiv AI
Sep 28

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation

MOPD‑Router rethinks teacher routing in multi‑teacher on‑policy distillation by routing supervision over the full teacher pool at each token, eliminating the need for prompt‑level domain labels or a separate routing model. The framework offers a plug‑in interface for various metrics, and introduces ExpertAlign, which scores teachers based on how well their corrections reflect their specialized post‑training knowledge. Experiments on both unlabeled and domain‑labeled mixtures show that ExpertAlign outperforms existing methods, improving overall scores by up to 12.3% on unlabeled data and 7.8% on domain‑labeled data. whyItMatters":"Token‑level routing enables the use of complementary supervision across domains without relying on domain labels, leading to significant performance gains in multi‑teacher distillation settings."

By Tianze Xu, Yanzhao Zheng, Zhentao Zhang, Yuanqiang Yu, Chao Ma, Jihuai Zhu, Lelun Wu, Lyumanshan Ye, Pengfei Liu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu
arXiv AI
Aug 18

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

arXiv:2608. 16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories.

By Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu
Hugging Face Trending Papers
Sep 28

PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

PMOPD introduces a projection-based approach to multi-teacher on-policy distillation, addressing the capability seesaw problem by constructing subspace memories from task-specific parameter displacements and projecting gradients and optimizer updates to avoid cross-task interference. It also includes a lightweight conflict probe for task interaction analysis, a task ordering strategy, and a cycling mechanism to balance subspace estimation and task revisitation. Experiments on Code, Reason, and Math tasks demonstrate that PMOPD improves all evaluated capabilities, raising average scores by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B.

arXiv AI
Sep 4

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

The paper introduces Teacher-Gated On-Policy Distillation (TGOPD), a method that verifies teacher reliability at the prompt level before applying dense supervision in on-policy distillation. TGOPD uses verifier-scored teacher probes to decide whether to route a prompt to dense OPD or to a verifier-grounded alternative. Experiments on 4B and 35B models across mathematics, code, and instruction tasks show TGOPD outperforms vanilla OPD and improves teacher GPU utilization from 9.8% to 78.9% in a 4B single-domain run.

By Zhiwei Zhang, Zechen Sun, Fei Zhao, Kang Peng, Bin Liang, Huayu Deng, Yao Hu, Kam-Fai Wong, Mu Chuan
arXiv AI
Aug 20

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

Open-MOPD addresses the capability imbalance problem in multi-teacher on-policy distillation (M-OPD) by isolating capability integration from routing ambiguity and revealing a 35.6% headroom gap compared to a domain-routed oracle ensemble. The study identifies three key factors—sequence-length disparities, convergence drift, and reward staleness—that misallocate token-level optimization budgets, leading to severe degradation in concise tasks. The proposed Open-MOPD framework introduces token-share balancing, gap-aware dynamic budget allocation, and student reward refresh, boosting headroom recovery to 83.4% and providing an open-source, reproducible post‑training recipe and evaluation suite.

By Huan-ang Gao, Haohan Chi, Yong Yan, Shiyuan Feng, Hanlin Wu, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou