Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
arXiv:2605. 07804v3 Announce Type: replace-cross Abstract: On-policy distillation (OPD) leverages dense teacher rewards to enhance reasoning models.
arXiv:2606. 21994v2 Announce Type: replace Abstract: On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories.
arXiv:2605. 07804v3 Announce Type: replace-cross Abstract: On-policy distillation (OPD) leverages dense teacher rewards to enhance reasoning models.
arXiv:2609.33455v2 Announce Type: replace Abstract: On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since e...
The paper introduces SCOUT, a co‑training framework that adapts an off‑policy teacher to better continue from student‑generated prefixes in on‑policy distillation (OPD). By periodically optimizing the teacher’s conditional continuation ability using reinforcement learning with verifiable rewards, SCOUT improves the teacher’s performance on student prefixes. Experiments across various teacher‑student setups, model scales, and reasoning domains show that SCOUT consistently enhances the effectiveness of OPD.
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conve...
arXiv:2608. 16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories.
arXiv:2609.36246v1 Announce Type: new Abstract: We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressivel...
arXiv:2609.37500v1 Announce Type: new Abstract: On-policy distillation (OPD) trains language models using dense token-level teacher supervision on student-generated trajectories. However, its relianc...
Offline on-policy distillation gains efficiency by collecting student trajectories and teacher supervision once and reusing them throughout optimization. The same reuse makes imperfect supervision per...
The paper introduces a method for offline on‑policy distillation that addresses the problem of imperfect teacher supervision. By training on teacher‑successful problems and measuring changes in token likelihoods on teacher‑failed trajectories, the authors derive a learnability signal that weights the distillation loss. This approach improves performance on mathematical reasoning and code generation tasks while reducing computational cost compared to online distillation.
arXiv:2609.39687v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises studen...
arXiv:2606. 08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts.
arXiv:2607. 26057v1 Announce Type: cross Abstract: On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected continuations that elicit unreliable supervision and waste compute.