arXiv:2608. 09745v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, providing a useful complement to reinforcement learning with sparse outcome rewards.
By Zhuo Sun, Entong Li, Yanlong Zhao, Xiaoyuan Cheng, Wenxuan Yuan, Kaiyu Li, Che Liu, Huihang Liu, Harrison Bo Hua Zhu, Li Zeng
arXiv:2609.37132v1 Announce Type: new
Abstract: On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level super...
By Zheng Zhang, Xinyue Tan, Lufei Li, Xinyi Zhang, Yexin Li, Kan Ren
The paper introduces a recursive self-improvement framework for language models that replaces an external teacher with a frozen copy of the student, enabling dynamic co-evolution (DCE) and self-refined concise learning (SRCL). DCE allows the privileged teacher to evolve alongside the student, while SRCL trains on shorter, verified rewrites to reduce verbosity. Experiments show that the combined DCE+SRCL approach outperforms traditional on‑policy self‑distillation across multiple model sizes and math benchmarks, achieving significant accuracy gains and shorter outputs.
By Shangjian Yin, Zehao Zhao, Kavosh Asadi, Rui Liu, Yuchen Lu, Shike Mei, Hang Cui, Luke Simon, Zhouxing Shi, Hamed Firooz
arXiv:2607. 02502v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access.
By Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Shuang Qiu, Linqi Song
RISE (Recursive Improvement via Self-Extrapolating Policy Distillation) is a new method that builds a synthetic teacher from a language model’s own RLVR training trajectory. By extrapolating the displacement between the current checkpoint and a trailing anchor in parameter or logit space, RISE transforms sparse outcome-based updates into dense token-level targets without external models or privileged conditioning. The approach recursively refines the student model, combining RLVR and on‑policy distillation, and demonstrates superior performance across mathematical reasoning, STEM, code generation, and multi‑turn agentic tasks.
By Yang Li, Semih Yavuz, Shafiq Joty
arXiv:2607. 18955v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation.
By Qiye Cai, Yichuan Ma, Linyang Li, Peiji Li, Yongkang Chen, Qipeng Guo, Yicheng Zou, Tao Gui, Xiaocheng Feng, Bing Qin
arXiv:2606. 04036v1 Announce Type: new Abstract: On-policy self-distillation, where a language model conditions on privileged context to supervise its own generations, is a promising source of dense supervision for sparse-reward reinforcement learning.
By Yifeng Liu, Shiyuan Zhang, Yifan Zhang, Quanquan Gu
arXiv:2609.08798v1 Announce Type: new
Abstract: Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important...
By Youngrok Park, Sangmin Bae, Hojung Jung, Jongwoo Ko, Yunseon Choi, Young Jin Kim, Pashmina Cameron, Aaron Courville, Se-Young Yun
arXiv:2607. 01763v1 Announce Type: new Abstract: Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities.
By Meng Wang, Haohan Zhao, Wenzhuo Liu, Lu Yang, Geng Liu, Haiyang Guo, Guo-Sen Xie, Gaofeng Meng, Hongbin Liu, Fei Zhu
arXiv:2609.24646v1 Announce Type: new
Abstract: On-policy self-distillation fine-tuning (SDFT) learns new skills from demonstrations while reducing forgetting, but it always distils toward the full d...
By Ahmed Khaled Khamis, Xiaotong Ji, Hassan Jaber, Rasul Tutunov, Matthieu Zimmer, Jun Wang, Haitham Bou-Ammar
ComputerSD is an online self‑distillation method for computer‑use agents that leverages real‑time feedback from executed GUI transitions. It uses a fine‑tuned GUI analyzer to generate guidance and a step‑level value score after each action, combining token‑level OPSD with trajectory‑level GRPO in an asynchronous training framework. On the OSWorld‑Verified benchmark, ComputerSD improves performance over outcome‑only GRPO by 1.9 and 4.1 percentage points on Qwen3‑VL‑8B‑Thinking and EvoCUA‑8B backbones, and shows strong generalizability in out‑of‑distribution tests.
By Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen
arXiv:2606. 09091v1 Announce Type: new Abstract: On-policy distillation (OPD) has recently emerged as an important post-training paradigm.
By Dongze Hao, Zhiwei Jin, Chen Chen, Haonan Lu