The paper investigates self‑distillation techniques for language models by systematically varying three key design choices: the source of rollout tokens (student vs. teacher), the teacher coupling strategy (frozen or exponential moving average), and the KL divergence direction (reverse or forward). Experiments on Qwen2.5‑7B and Ministral‑3‑3B across 1,200 adaptation runs reveal that rollout source mainly affects acquisition on contradictory tasks, teacher coupling most strongly influences acquisition across all tasks, and KL direction impacts retention differently depending on the model. A controlled theoretical model reproduces these empirical trends, offering a unified framework for understanding acquisition‑retention trade‑offs in self‑distillation.
By Luis Zuin, Alexis Huet, Dario Rossi, Zied Ben Houidi
arXiv:2605.10194v2 Announce Type: replace
Abstract: On-policy self-distillation (OPSD) uses a model as its own teacher under privileged context, providing token-level supervision on the model's own r...
By Jiaxuan Wang, Xuan Ouyang, Zhiyu Chen, Yulan Hu, Lan-Zhe Guo
The paper investigates how on-policy self‑distillation can alter a model’s behavior by conditioning on privileged information. It contrasts attractive self‑distillation, which pulls a model toward a privileged teacher, with repulsive self‑distillation, which pushes it away, showing that attraction reduces exploratory reasoning while repulsion lengthens responses and can destabilize the model. The authors propose a contrastive self‑distillation objective that combines attraction to a correct‑solution teacher with repulsion from an incorrect‑solution teacher, finding that this approach improves reasoning performance across various model types while keeping response lengths stable.
By Anton Baumann, Akmal Ashirmatov, Leo Schmidt-Traub, Frederike L\"ubeck, Jonas H\"ubotter, Thomas Kleine Buening, Andreas Krause
arXiv:2609.37041v1 Announce Type: cross
Abstract: Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an im...
By Su Ee Tan, Xiaotong Ji, Rasul Tutunov, Haitham Bou-Ammar, Matthieu Zimmer
The paper investigates on‑policy self‑distillation (OPSD), where a student model learns from its own outputs using token‑level supervision conditioned on privileged reference information. Experiments with Qwen3 models on science and mathematics datasets show that the correct reference does not consistently improve performance; students can improve without it, and solutions from other problems sometimes outperform the correct reference. The study finds that student predictions align more closely with the base model’s reasoning than with the reference supervision, and that alignment alone does not reliably predict performance gains.
By Samyak Shrestha, Alexander Tessier
arXiv:2609.38342v1 Announce Type: new
Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning....
By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu