arXiv AI
Sep 7

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation

RISE (Recursive Improvement via Self-Extrapolating Policy Distillation) is a new method that builds a synthetic teacher from a language model’s own RLVR training trajectory. By extrapolating the displacement between the current checkpoint and a trailing anchor in parameter or logit space, RISE transforms sparse outcome-based updates into dense token-level targets without external models or privileged conditioning. The approach recursively refines the student model, combining RLVR and on‑policy distillation, and demonstrates superior performance across mathematical reasoning, STEM, code generation, and multi‑turn agentic tasks.

By Yang Li, Semih Yavuz, Shafiq Joty
arXiv Computation and Language
6d ago

Recursive Self-Improvement via On-Policy Distillation for Reasoning

The paper introduces a recursive self-improvement framework for language models that replaces an external teacher with a frozen copy of the student, enabling dynamic co-evolution (DCE) and self-refined concise learning (SRCL). DCE allows the privileged teacher to evolve alongside the student, while SRCL trains on shorter, verified rewrites to reduce verbosity. Experiments show that the combined DCE+SRCL approach outperforms traditional on‑policy self‑distillation across multiple model sizes and math benchmarks, achieving significant accuracy gains and shorter outputs.

By Shangjian Yin, Zehao Zhao, Kavosh Asadi, Rui Liu, Yuchen Lu, Shike Mei, Hang Cui, Luke Simon, Zhouxing Shi, Hamed Firooz
arXiv AI
Jul 3

DemoPSD: Disagreement-Modulated Policy Self-Distillation

arXiv:2607. 02502v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access.

By Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Shuang Qiu, Linqi Song
arXiv Machine Learning
2d ago

Activation-Conditioned Self-Distillation

arXiv:2609.38342v1 Announce Type: new Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning....

By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu
arXiv AI
Jun 9

Trajectory-Refined Distillation

arXiv:2606. 08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts.

By Li Jiang, Haoran Xu, Yichuan Ding, Amy Zhang