The paper introduces a recursive self-improvement framework for language models that replaces an external teacher with a frozen copy of the student, enabling dynamic co-evolution (DCE) and self-refined concise learning (SRCL). DCE allows the privileged teacher to evolve alongside the student, while SRCL trains on shorter, verified rewrites to reduce verbosity. Experiments show that the combined DCE+SRCL approach outperforms traditional on‑policy self‑distillation across multiple model sizes and math benchmarks, achieving significant accuracy gains and shorter outputs.
By Shangjian Yin, Zehao Zhao, Kavosh Asadi, Rui Liu, Yuchen Lu, Shike Mei, Hang Cui, Luke Simon, Zhouxing Shi, Hamed Firooz
arXiv:2609.33455v2 Announce Type: replace
Abstract: On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since e...
By Zizhuo Lin, Quanling Liu, Yi Yang, Yawei Luo
Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation explores how on‑policy distillation (OPD) can be improved by addressing teacher uncertainty that contracts as the teacher continues from a student‑generated prefix. The authors identify Teacher Uncertainty Contraction (TUC) and theoretically analyze its variance‑bias trade‑off, leading to the proposal of Adaptive‑Continuations On‑Policy Distillation (AC‑OPD). Experiments on mathematical reasoning and code generation show that AC‑OPD consistently outperforms standard OPD, with controlled‑continuation and matched‑budget analyses supporting the adaptive‑continuation design.
By Jingang Zhou, Yuyi Zhou, Haiyang Guo, Xukai Wang, Shuai Feng, Sirui Gao, Jian Xu, Qingpei Guo, Xu-Yao Zhang
arXiv:2608. 19408v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing student-generated trajectories with dense token-level supervision from a teacher.
By Chen Yang, Haiyuan Wan, Rengrong Xiong, Yize Chen, Danny H. K. Tsang
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conve...
arXiv:2606. 21994v2 Announce Type: replace Abstract: On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories.
By Qingfei Zhao, Huan Song, Shuyu Tian, Jiawei Shao, Xuelong Li