arXiv Machine Learning By Jingang Zhou, Yuyi Zhou, Haiyang Guo, Xukai Wang, Shuai Feng, Sirui Gao, Jian Xu, Qingpei Guo, Xu-Yao Zhang

Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation

Read the original on arXiv Machine Learning →

Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation explores how on‑policy distillation (OPD) can be improved by addressing teacher uncertainty that contracts as the teacher continues from a student‑generated prefix. The authors identify Teacher Uncertainty Contraction (TUC) and theoretically analyze its variance‑bias trade‑off, leading to the proposal of Adaptive‑Continuations On‑Policy Distillation (AC‑OPD). Experiments on mathematical reasoning and code generation show that AC‑OPD consistently outperforms standard OPD, with controlled‑continuation and matched‑budget analyses supporting the adaptive‑continuation design.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 9

Trajectory-Refined Distillation

arXiv:2606. 08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts.

By Li Jiang, Haoran Xu, Yichuan Ding, Amy Zhang
arXiv AI
Sep 17

Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers

The paper introduces a method for offline on‑policy distillation that addresses the problem of imperfect teacher supervision. By training on teacher‑successful problems and measuring changes in token likelihoods on teacher‑failed trajectories, the authors derive a learnability signal that weights the distillation loss. This approach improves performance on mathematical reasoning and code generation tasks while reducing computational cost compared to online distillation.

By Yihao Ai, Weilong Yan