arXiv:2609.23697v1 Announce Type: cross
Abstract: Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, howe...
By Jie Sun, Mao Zheng, Mingyang Song, Zeyuan Liu, Gengsheng Li, Houcheng Jiang, Yilin Cheng, Bichuan Feng, Yuchen Cai, Junfeng Fang, Xiang Wang
arXiv:2607. 16246v1 Announce Type: cross Abstract: Off-policy distillation is now central to large language model pre-training, yet how training data, objective parameterization, and model capabilities interact remains poorly characterized.
By Jiangan Yuan, Zhixuan Li, Han Xu
arXiv:2608. 16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories.
By Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu
The paper investigates on‑policy distillation (OPD), showing that teacher supervision during OPD contains significant noise that grows with teacher size, yet the student policy remains largely unaffected by this noise. It finds that OPD’s gains stem mainly from suppressing low‑log‑probability tokens, a process that can be replicated without a teacher. Building on this insight, the authors propose On‑Policy Self‑Adaptation (OPSA), a supervision‑free method that uses entropy‑adaptive negative advantages to improve performance on several benchmarks, outperforming both the base model and OPD.
By Yi Ding, Ruqi Zhang
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conve...
The paper introduces uncertainty‑calibrated Multi‑Teacher On‑Policy Distillation (MOPD) to better preserve general language model capabilities while specializing to specific domains. By employing dual‑temperature sampling, positive‑advantage‑density filtering, and centered log‑likelihood filtering, the method selects more informative trajectories and token updates, leading to significant improvements in general‑capability performance on role‑playing and medical‑domain tasks without sacrificing domain performance.
By Ziyuan Liu, Jiao Ou, Jian Liang, Ruiming Tang, Cheng Luo