arXiv:2609.25623v1 Announce Type: new
Abstract: More privileged information does not always make a better teacher. We study this tension in on-policy self-distillation (OPSD), where a frozen copy of...
By Kanghui Tian, Siyuan Liu, Tianxiang Jiang, Shuai Dong, Yizhuo Li, Tian Ding, Yuan Guo, Songze Li, Haowen Hou, Congcong Wang, Yi Wang
arXiv:2609.37915v1 Announce Type: new
Abstract: On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies...
By Md. Ismail Hossain, Humaira Kousar, Isidora Chara Tourni
The paper investigates how privileged information—such as a teacher’s full solution or reasoning trace—affects on‑policy self‑distillation (OPSD) in language models. Using the AMPLE‑Math benchmark, the authors compare distillation with and without extra teacher views, finding that reference‑free distillation explains most gains for Qwen3‑1.7B, while additional references provide modest benefits, especially for polished solutions. The study also shows that the impact of privileged data depends on the student’s training regime and that altering token‑level supervision can leave student behavior largely unchanged.
By XiuYu Zhang, Wei Chow, Junfeng Fang, Zhenkai Liang, Tat-Seng Chua
The paper investigates on‑policy self‑distillation (OPSD), where a student model learns from its own outputs using token‑level supervision conditioned on privileged reference information. Experiments with Qwen3 models on science and mathematics datasets show that the correct reference does not consistently improve performance; students can improve without it, and solutions from other problems sometimes outperform the correct reference. The study finds that student predictions align more closely with the base model’s reasoning than with the reference supervision, and that alignment alone does not reliably predict performance gains.
By Samyak Shrestha, Alexander Tessier
arXiv:2609.21619v1 Announce Type: new
Abstract: On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. How...
By Qiangqiang He, Jin Li, MingCai Chen
arXiv:2608. 01735v2 Announce Type: replace Abstract: On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training.
By Jianyu Wu, Yizhou Wang, Encheng Su, Chen Tang, Shixiang Tang
On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of...
The paper reviews On‑Policy Self‑Distillation (OPSD), a method where a language model learns from its own generations using privileged information such as reference solutions or plans, eliminating the need for a larger teacher model. It identifies a key failure mode—collapse, where the model’s reasoning paths narrow progressively—and analyzes it through three levers: signal application, privileged information, and teacher dynamics. The review focuses on mathematical reasoning, offering a unified vocabulary and distinguishing settled facts from ongoing debates.
By Justin Robert, Raheel Qader
arXiv:2608. 04794v1 Announce Type: new Abstract: Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, conditioned on privileged information (PI) about the answer such as a reference solution, supplies dense per-token supervision to a student that never sees it.
By Sarthak Harne, Chinmay Karkar, Yash Pandya, Ahmed Awadallah, Akshay Nambi
The paper introduces On-Policy Attention Self-Distillation (OPASD), a method that augments token-level supervision with solution-conditioned attention distillation for reasoning models. OPASD projects a privileged teacher’s attention onto student-visible positions, renormalizes the distribution, and aligns it with the student. Experiments on three model sizes and four math benchmarks show that OPASD improves accuracy by 4.98–8.40 percentage points, reduces generated tokens by 73.9%, cuts compute by 72.6%, and trains 1.53× faster compared to token-only distillation.
By Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil, Md. Faiyaz Abdullah Sayeedi, Tasnim Mohiuddin, Md Mofijul Islam
arXiv:2607. 26246v1 Announce Type: new Abstract: On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for transferring capabilities across LLMs.
By Fangxu Yu, Zinan Lin, Xiaodong Liu, Weijia Xu, Michael Xu, Tianyi Zhou, Jianfeng Gao
arXiv:2609.37132v1 Announce Type: new
Abstract: On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level super...
By Zheng Zhang, Xinyue Tan, Lufei Li, Xinyi Zhang, Yexin Li, Kan Ren