arXiv:2609.36546v1 Announce Type: cross
Abstract: On-policy distillation (OPD) trains a student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD ma...
By Shutong Wu, Xiwen Chen, Brendan Rappazzo, Daiheng Zhang, Anderson Schneider, Yuriy Nevmyvaka, Jiawei Zhang
arXiv:2606. 06021v1 Announce Type: new Abstract: On-policy distillation (OPD) supervises the student only in output space by matching next-token probabilities.
By Shenzhi Yang, Guangcheng Zhu, Bowen Song, Haobo Wang, Mingxuan Xia, Xing Zheng, Yingfan Ma, Zhongqi Chen, Weiqiang Wang, Gang Chen
The paper introduces Selective Supervision for Direct-OPD (S$^2$D-OPD), a refinement of Direct On-Policy Distillation that filters out states where the teacher’s policy change is minimal, as measured by the teacher‑reference Jensen‑Shannon divergence. By masking low‑divergence states and keeping only the top 10% of states per response, S$^2$D-OPD improves held‑out accuracy on AIME and HMMT benchmarks across multiple teacher‑student pairs without additional forward passes.
By Yibo Zhao, Zixuan Yang, Yunshi Lan, Xiang Li
arXiv:2609.37915v1 Announce Type: new
Abstract: On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies...
By Md. Ismail Hossain, Humaira Kousar, Isidora Chara Tourni
arXiv:2608. 09745v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, providing a useful complement to reinforcement learning with sparse outcome rewards.
By Zhuo Sun, Entong Li, Yanlong Zhao, Xiaoyuan Cheng, Wenxuan Yuan, Kaiyu Li, Che Liu, Huihang Liu, Harrison Bo Hua Zhu, Li Zeng
arXiv:2606. 10385v1 Announce Type: cross Abstract: On-policy distillation (OPD) has demonstrated strong empirical gains in enhancing complex reasoning in LLMs by aligning a student model with a teacher's predictive distribution over the student's own trajectories.
By Wenhao Zhang
arXiv:2608. 09447v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state mismatch of offline distillation.
By Zehao Chen, Gongxun Li, Tianxiang Ai, Yifei Li, Zixuan Huang, Wang Zhou, Tao Huang, Fuzhen Zhuang, Xianglong Liu, Jianxin Li, Deqing Wang, Yikun Ban
arXiv:2605.10194v2 Announce Type: replace
Abstract: On-policy self-distillation (OPSD) uses a model as its own teacher under privileged context, providing token-level supervision on the model's own r...
By Jiaxuan Wang, Xuan Ouyang, Zhiyu Chen, Yulan Hu, Lan-Zhe Guo
arXiv:2608. 08726v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) uses a privileged teacher to supervise a reasoning model on prefixes sampled from its own rollouts.
By Yangyang Feng, Zhuoyan Feng, Junlan Chen
arXiv:2608. 05219v1 Announce Type: new Abstract: Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access to training-only references, such as successful trajectories.
By Junzhuo Liu, Weiwei Li, Jun Ling, Peng Wang
Self-OPD introduces a teacher‑free on‑policy distillation framework for flow matching models, using the student’s own exploration to generate step‑wise supervision. At each timestep the deterministic next‑state prediction is branched into multiple stochastic SDE candidates, rolled out, and compared against a deterministic baseline to compute normalized advantages. The velocity field is then optimized with a pull‑push objective that attracts high‑advantage branches and repels low‑advantage ones, while multi‑objective alignment is achieved by fusing normalized scores at the reward level.
By Shiyi Zhang, Mushui Liu, Yunze Tong, Wanggui He, Siyu Zou, Jinlong Liu, Yunlong Yu, Jian Song, Hao Jiang, Pipei Huang, Bo Zheng
The paper introduces On-Policy Attention Self-Distillation (OPASD), a method that augments token-level supervision with solution-conditioned attention distillation for reasoning models. OPASD projects a privileged teacher’s attention onto student-visible positions, renormalizes the distribution, and aligns it with the student. Experiments on three model sizes and four math benchmarks show that OPASD improves accuracy by 4.98–8.40 percentage points, reduces generated tokens by 73.9%, cuts compute by 72.6%, and trains 1.53× faster compared to token-only distillation.
By Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil, Md. Faiyaz Abdullah Sayeedi, Tasnim Mohiuddin, Md Mofijul Islam