arXiv AI

The Role of Feedback Alignment in Self-Distillation

arXiv:2606. 11173v1 Announce Type: new Abstract: Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response.

arXiv AI
3d ago

Disentangling Self-Distillation: Measuring and Modeling Acquisition and Retention

The paper investigates self‑distillation techniques for language models by systematically varying three key design choices: the source of rollout tokens (student vs. teacher), the teacher coupling strategy (frozen or exponential moving average), and the KL divergence direction (reverse or forward). Experiments on Qwen2.5‑7B and Ministral‑3‑3B across 1,200 adaptation runs reveal that rollout source mainly affects acquisition on contradictory tasks, teacher coupling most strongly influences acquisition across all tasks, and KL direction impacts retention differently depending on the model. A controlled theoretical model reproduces these empirical trends, offering a unified framework for understanding acquisition‑retention trade‑offs in self‑distillation.

By Luis Zuin, Alexis Huet, Dario Rossi, Zied Ben Houidi
arXiv AI
Sep 21

On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation

The paper investigates how on-policy self‑distillation can alter a model’s behavior by conditioning on privileged information. It contrasts attractive self‑distillation, which pulls a model toward a privileged teacher, with repulsive self‑distillation, which pushes it away, showing that attraction reduces exploratory reasoning while repulsion lengthens responses and can destabilize the model. The authors propose a contrastive self‑distillation objective that combines attraction to a correct‑solution teacher with repulsion from an incorrect‑solution teacher, finding that this approach improves reasoning performance across various model types while keeping response lengths stable.

By Anton Baumann, Akmal Ashirmatov, Leo Schmidt-Traub, Frederike L\"ubeck, Jonas H\"ubotter, Thomas Kleine Buening, Andreas Krause
arXiv Machine Learning
Aug 20

Rethinking Privileged Information in On-Policy Self-Distillation

The paper investigates on‑policy self‑distillation (OPSD), where a student model learns from its own outputs using token‑level supervision conditioned on privileged reference information. Experiments with Qwen3 models on science and mathematics datasets show that the correct reference does not consistently improve performance; students can improve without it, and solutions from other problems sometimes outperform the correct reference. The study finds that student predictions align more closely with the base model’s reasoning than with the reference supervision, and that alignment alone does not reliably predict performance gains.

By Samyak Shrestha, Alexander Tessier
arXiv Machine Learning
2d ago

Activation-Conditioned Self-Distillation

arXiv:2609.38342v1 Announce Type: new Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning....

By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu
arXiv AI
Aug 6

Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation

arXiv:2608. 04794v1 Announce Type: new Abstract: Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, conditioned on privileged information (PI) about the answer such as a reference solution, supplies dense per-token supervision to a student that never sees it.

By Sarthak Harne, Chinmay Karkar, Yash Pandya, Ahmed Awadallah, Akshay Nambi
arXiv Machine Learning
Sep 24

RL Starts before RL: On Policy Distillation for Better Reinforcement Learning

The paper investigates on‑policy distillation (OPD) as a preparatory step for reinforcement learning (RL). It shows that students initialized with OPD achieve higher final RL performance than those trained directly with RL or with supervised fine‑tuning followed by RL, even when OPD offers little immediate accuracy gain. The study also finds that the choice of distillation objective (reverse‑KL vs forward‑KL) and the source of trajectories influence OPD’s effectiveness at different stages of RL training.

By Shuai Dong, Yongfu Zhu, Yuqi Xu, Weichu Xie, Liuwenpu, Ziyue Wang, Kaiwen Tuo, Congcong Wang, Siyuan Wang, Wenqi Shao, Shuai Yang, Ji Zhao, Caoyuan Ma, Wenzheng Chang, Taiqiang Wu, Xinlei Yu, Hongrui Wu, Xiaoxuan He, Fangke Chen, Dianyi Wang, Kanghui Tian, Sirry Chen, Xingyu Liu, Xiangnan Wu, Jiawei Guo, Haowen Hou, LingHan Chen, Zhongyu Wei, Jiaqi Wang