When Clipping Reverses Correction: Failure Dynamics of Pointwise Forward-KL On-Policy Self-Distillation
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2606. 08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts.
The paper investigates self‑distillation techniques for language models by systematically varying three key design choices: the source of rollout tokens (student vs. teacher), the teacher coupling strategy (frozen or exponential moving average), and the KL divergence direction (reverse or forward). Experiments on Qwen2.5‑7B and Ministral‑3‑3B across 1,200 adaptation runs reveal that rollout source mainly affects acquisition on contradictory tasks, teacher coupling most strongly influences acquisition across all tasks, and KL direction impacts retention differently depending on the model. A controlled theoretical model reproduces these empirical trends, offering a unified framework for understanding acquisition‑retention trade‑offs in self‑distillation.
arXiv:2607. 02502v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access.
arXiv:2609.32201v2 Announce Type: replace Abstract: On-Policy Context Distillation (OPCD) has recently emerged as a powerful technique for transferring context to student models and for self-improvem...
The paper reviews On‑Policy Self‑Distillation (OPSD), a method where a language model learns from its own generations using privileged information such as reference solutions or plans, eliminating the need for a larger teacher model. It identifies a key failure mode—collapse, where the model’s reasoning paths narrow progressively—and analyzes it through three levers: signal application, privileged information, and teacher dynamics. The review focuses on mathematical reasoning, offering a unified vocabulary and distinguishing settled facts from ongoing debates.
arXiv:2607. 05184v1 Announce Type: new Abstract: Self-distillation is a promising recipe for self-improvement in language models.