arXiv AI
Aug 20

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

Open-MOPD addresses the capability imbalance problem in multi-teacher on-policy distillation (M-OPD) by isolating capability integration from routing ambiguity and revealing a 35.6% headroom gap compared to a domain-routed oracle ensemble. The study identifies three key factors—sequence-length disparities, convergence drift, and reward staleness—that misallocate token-level optimization budgets, leading to severe degradation in concise tasks. The proposed Open-MOPD framework introduces token-share balancing, gap-aware dynamic budget allocation, and student reward refresh, boosting headroom recovery to 83.4% and providing an open-source, reproducible post‑training recipe and evaluation suite.

By Huan-ang Gao, Haohan Chi, Yong Yan, Shiyuan Feng, Hanlin Wu, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
arXiv AI
3d ago

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

arXiv:2609.37898v1 Announce Type: new Abstract: Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-st...

By Youling Huang, Tiankuo Xu, Jiaji Liu, Tong Zheng, Shuo Zhou, Shaotong Qi, Junchi Yao, Shiyang Liu, Hao Xu, Pengcheng Xu, Bo Huang, Hongyi Fu, Lin Lin