CLOOPD: Closing the Learner Loop in On-Policy Distillation
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2609.14636v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a standard approach for transferring capabilities from large teachers to compact students. Its cost, however, i...
Open-MOPD addresses the capability imbalance problem in multi-teacher on-policy distillation (M-OPD) by isolating capability integration from routing ambiguity and revealing a 35.6% headroom gap compared to a domain-routed oracle ensemble. The study identifies three key factors—sequence-length disparities, convergence drift, and reward staleness—that misallocate token-level optimization budgets, leading to severe degradation in concise tasks. The proposed Open-MOPD framework introduces token-share balancing, gap-aware dynamic budget allocation, and student reward refresh, boosting headroom recovery to 83.4% and providing an open-source, reproducible post‑training recipe and evaluation suite.
arXiv:2609.36546v1 Announce Type: cross Abstract: On-policy distillation (OPD) trains a student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD ma...
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conve...
arXiv:2609.37898v1 Announce Type: new Abstract: Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-st...
arXiv:2609.37170v1 Announce Type: cross Abstract: Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation tr...