arXiv Machine Learning By Nicolas Zucchet, Scott W. Linderman

Divergence controls entropy in distillation

Read the original on arXiv Machine Learning →

The paper investigates how the choice of divergence in knowledge distillation affects the entropy of the student model. It shows that forward KL increases student entropy beyond the teacher’s, while reverse KL decreases it, and that interpolating between the two yields a smooth entropy change early in training but a sharp shift at convergence. The study also finds that on‑policy distillation’s lower entropy stems from token‑level reverse KL rather than sampling, positioning divergence as an implicit entropy regularizer, especially evident in self‑distillation scenarios.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 2

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

arXiv:2609.01532v1 Announce Type: new Abstract: Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits...

By Jacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li, Hanqing Zeng, Yinglong Xia, Benyu Zhang, Zhuokai Zhao, Qiang Zhang, Pang Wei Koh, Luke Zettlemoyer, Wen-tau Yih
arXiv Machine Learning
Jun 4

Self-Distilled Policy Gradient

arXiv:2606. 04036v1 Announce Type: new Abstract: On-policy self-distillation, where a language model conditions on privileged context to supervise its own generations, is a promising source of dense supervision for sparse-reward reinforcement learning.

By Yifeng Liu, Shiyuan Zhang, Yifan Zhang, Quanquan Gu