OPSRD: On-Policy Self-Role Distillation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.37132v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level super...
RISE (Recursive Improvement via Self-Extrapolating Policy Distillation) is a new method that builds a synthetic teacher from a language model’s own RLVR training trajectory. By extrapolating the displacement between the current checkpoint and a trailing anchor in parameter or logit space, RISE transforms sparse outcome-based updates into dense token-level targets without external models or privileged conditioning. The approach recursively refines the student model, combining RLVR and on‑policy distillation, and demonstrates superior performance across mathematical reasoning, STEM, code generation, and multi‑turn agentic tasks.
The paper introduces a recursive self-improvement framework for language models that replaces an external teacher with a frozen copy of the student, enabling dynamic co-evolution (DCE) and self-refined concise learning (SRCL). DCE allows the privileged teacher to evolve alongside the student, while SRCL trains on shorter, verified rewrites to reduce verbosity. Experiments show that the combined DCE+SRCL approach outperforms traditional on‑policy self‑distillation across multiple model sizes and math benchmarks, achieving significant accuracy gains and shorter outputs.
arXiv:2607. 02502v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access.
arXiv:2609.38342v1 Announce Type: new Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning....
arXiv:2606. 08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts.