arXiv:2603. 07079v3 Announce Type: replace Abstract: On-policy distillation is a promising approach for transferring knowledge between language models, where a student learns from dense token-level signals along its own trajectories.
By Woogyeol Jin, Taywon Min, Yongjin Yang, Dennis Wei, Yi Zhou, Swanand Ravindra Kadhe, Nathalie Baracaldo, Kimin Lee
arXiv:2608.29846v1 Announce Type: cross
Abstract: Sampled-token on-policy distillation (OPD) efficiently transfers capabilities from teacher to student using student-generated tokens, requiring teach...
By Run Yang, Runpeng Dai, Jie Sun, Jielei Zhang, Fan Zhou, Hongtu Zhu, Peiyi Li, Longwen Gao
arXiv:2608. 14685v1 Announce Type: new Abstract: Knowledge distillation (KD) is widely used to transfer the capabilities of large language models (LLMs) to smaller students, but existing objectives often struggle to balance faithful imitation and robust generation.
By Shizhen Li, Zhiyu Shen, Yuyin Lu, Yunhe Pang, Jielin Song, Yanghui Rao, Fu Lee Wang
arXiv:2609.01532v1 Announce Type: new
Abstract: Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits...
By Jacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li, Hanqing Zeng, Yinglong Xia, Benyu Zhang, Zhuokai Zhao, Qiang Zhang, Pang Wei Koh, Luke Zettlemoyer, Wen-tau Yih
arXiv:2609.36734v1 Announce Type: new
Abstract: Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implic...
By Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty
arXiv:2606. 04036v1 Announce Type: new Abstract: On-policy self-distillation, where a language model conditions on privileged context to supervise its own generations, is a promising source of dense supervision for sparse-reward reinforcement learning.
By Yifeng Liu, Shiyuan Zhang, Yifan Zhang, Quanquan Gu
The paper introduces Tail‑Corrected Top‑k On‑Policy Distillation (TT‑OPD), a method that improves on existing Top‑k OPD by combining the selected top‑k tokens with a sampled token from the student’s rollout. This hybrid approach recovers the probability mass discarded by limiting to top‑k, yielding an unbiased estimator of the reverse KL divergence gradient while maintaining low computational cost. Experiments show TT‑OPD outperforms other OPD variants in accuracy.
By Linjian Meng, Siyuan Gan, YuHan Li, Xiran Wang, Ziyang Ding, Ditang Gou, Yiming Wu, Zhen Zhao
arXiv:2607. 02502v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access.
By Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Shuang Qiu, Linqi Song
arXiv:2609.36246v1 Announce Type: new
Abstract: We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressivel...
By Haojin Wang, Dylan Zhang, Huaibo Chen, Suhao Yu, Yihang Sun, Zhanyang Jin, Jiaying Ye, Dianqi Li, Prasanna Sattigeri, Kamal Youcef-Toumi, Hao Peng
arXiv:2608.22898v1 Announce Type: new
Abstract: Diffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) large language models (LLMs), but their degraded gene...
By Hyeongsoo Lim, Jinyoung Kim, Eunseo Seo, Minho Jang, Jiwon Yoon
arXiv:2608. 14728v1 Announce Type: cross Abstract: On-policy distillation (OPD) has emerged as an effective paradigm for transferring knowledge between language models, where a student is trained to align its next-token distribution with the teacher's along its own trajectories.
By Huipeng Huang, Hongxin Wei
arXiv:2609.36695v1 Announce Type: new
Abstract: Self-distillation turns knowledge distillation into a closed learning loop and offers a path toward recursive self-improvement. Without an external tea...
By Rui Wang, Ruijie Wang, Bo Chen, Jiangxuan Long, Yingyu Liang