arXiv:2605.26872v4 Announce Type: replace-cross
Abstract: LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations....
By Zhengyu Hu, Zheyuan Xiao, Linxin Song, Fengqing Jiang, Yuetai Li, Zhihan Xiong, Yue Liu, Junhao Lin, Yao Su, Lijie Hu, Kaize Ding, Teng Xiao, Radha Poovendran
Activation-Conditioned Self-Distillation (ACSD) is a new on‑policy self‑distillation method that uses a frozen copy of the base model to extract a steering vector by contrasting activations from self‑generated trajectories that reach verified correct answers with all other trajectories. The student learns from next‑token distributions on its own prefixes, without needing reference text or teacher parameter updates, and is used alone at inference. Across five models, ACSD achieves the highest mean accuracy on four mathematical benchmarks, with notable gains on DeepSeek‑R1‑0528‑Qwen3‑8B and LiveCodeBench v6 compared to the OPSD baseline.
By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu
arXiv:2610.02703v1 Announce Type: new
Abstract: On-policy distillation (OPD) improves large language model reasoning by training students on their own rollouts with dense token-wise supervision from...
By Yuxiang Zhang, Ding Cao, Shuting Cui, Lei Wang, Weijieying Ren, Tianxiang Zhao
On-policy self-distillation (OPSD) improves reasoning by training a problem-only student on its own rollouts using dense token-level supervision from a privileged teacher that also sees a reference so...
The paper introduces VISTA, a method that enhances on‑policy self‑distillation (OPSD) by adapting the teacher model toward the student’s distribution using outcome‑verified rollouts. VISTA keeps the standard OPSD student update but selectively adjusts the teacher only on the top‑k positions with the largest teacher‑student KL divergence, without adding new sampling or reward objectives. Experiments on AIME24, AIME25, and HMMT25 with Qwen3 models show that VISTA outperforms OPSD across all scales, improving Avg@12 by up to 2.1 points.
By Zewen Ding, Zezhong Wu, Zhou Tao, Shida Wang, Shizhuo Hou, YongXiang Hua, Haoyu Cao, Linli Xu
arXiv:2402. 14035v4 Announce Type: replace-cross Abstract: Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, and modality.
By Zichang Liu, Qingyun Liu, Yuening Li, Liang Liu, Anshumali Shrivastava, Shuchao Bi, Lichan Hong, Ed H. Chi, Zhe Zhao
arXiv:2609.38025v1 Announce Type: cross
Abstract: On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OP...
By Zhenyu Wang, Tianze Wang, Linjun Zhang, Yifan Hu
arXiv:2605.28791v2 Announce Type: replace-cross
Abstract: On-policy self-distillation (SD) improves LLM reasoning by using teacher-side privileged information (PI) to turn sparse verifier outcomes in...
By Jiazhen Huang, Xiao Chen, Xiao Luo, Yong Dai, Senkang Hu, Yuzhi Zhao
arXiv:2608. 19408v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing student-generated trajectories with dense token-level supervision from a teacher.
By Chen Yang, Haiyuan Wan, Rengrong Xiong, Yize Chen, Danny H. K. Tsang
The paper investigates data efficiency and selection in On‑Policy Distillation (OPD) for large language models. It shows that 1‑shot OPD—training on a single example—consistently improves performance, especially when the example is hard, and that longer chain‑of‑thought (CoT) paths drive the gains rather than token entropy. Based on these findings, the authors propose a simple hard‑example selection strategy that, using only eight carefully chosen hard examples, matches the performance of a 17,000‑example baseline across models from 1.5B to 7B parameters.
By Zhinan Hou, Jiaqi Zhang, Xunliang Cai, Keyou You
arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.
By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan
On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable.