The paper introduces a method for offline on‑policy distillation that addresses the problem of imperfect teacher supervision. By training on teacher‑successful problems and measuring changes in token likelihoods on teacher‑failed trajectories, the authors derive a learnability signal that weights the distillation loss. This approach improves performance on mathematical reasoning and code generation tasks while reducing computational cost compared to online distillation.
By Yihao Ai, Weilong Yan
Offline on-policy distillation gains efficiency by collecting student trajectories and teacher supervision once and reusing them throughout optimization. The same reuse makes imperfect supervision per...
arXiv:2606. 29287v1 Announce Type: new Abstract: Diffusion and continuous-flow generative models achieve high-quality generation, and their deterministic sampling can be formulated as solving learned ODE dynamics.
By Chen Wang, Peiran Yun, Pan Xie, Ke Deng
arXiv:2512. 19643v2 Announce Type: replace Abstract: Numerical simulation of time-dependent partial differential equations (PDEs) is central to scientific and engineering applications, but high-fidelity solvers are often prohibitively expensive for long-horizon or time-critical settings.
By Rajyasri Roy, Dibyajyoti Nayak, Somdatta Goswami
arXiv:2503.19081v2 Announce Type: replace
Abstract: Scientific foundation models (SciFMs) aim to learn generalizable representations of physical systems governed by partial differential equations (PD...
By Serge Kotchourko, Amin Totounferoush, Michael W. Mahoney, Steffen Staab
arXiv:2606. 09949v1 Announce Type: cross Abstract: Data-driven PDE surrogates are trained with data produced by numerical PDE solvers.
By Pierre Cesar (DATAMOVE), Sofya Dymchenko (DATAMOVE), Abhishek Purandare (DATAMOVE), Bruno Raffin (DATAMOVE)
arXiv:2608. 16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories.
By Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu
arXiv:2609.38342v1 Announce Type: new
Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning....
By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu
arXiv:2609.33455v2 Announce Type: replace
Abstract: On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since e...
By Zizhuo Lin, Quanling Liu, Yi Yang, Yawei Luo
arXiv:2607. 29494v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense teacher supervision along student-generated trajectories, but its online rollout process incurs substantial computational cost, particularly when a few long responses delay batch completion.
By Qian Tan, Huaifei Liang, Xuanyu Zhu, Lei Jiang, Yuqiang Li
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conve...
arXiv:2606. 21994v2 Announce Type: replace Abstract: On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories.
By Qingfei Zhao, Huan Song, Shuyu Tian, Jiawei Shao, Xuelong Li