arXiv:2609.22566v1 Announce Type: cross
Abstract: Knowledge distillation (KD) aims to compress high-performance teacher LLMs into lightweight students. However, distilled students often exhibit subst...
By Dileesha Kannangara, Sanghamitra Dutta
The paper studies weak-to-strong generalization (W2SG), where a student model trained on a weaker teacher’s labels surpasses the teacher on the target task. Using a Bregman divergence bias‑variance decomposition, it shows that the student‑teacher risk gap depends on their expected misfit, without requiring convexity of the student hypothesis class. For squared loss, a sufficient condition is that the student converges to the teacher’s posterior mean, achievable by enlarging the student; for cross‑entropy loss, reducing the student’s predictive entropy and using reverse cross‑entropy can promote W2SG, which is empirically validated.
By Gengze Xu, Wei Yao, Ziqiao Wang, Yong Liu
The paper investigates how to choose the best quantized model from a family of compressed versions when target labels are scarce or unavailable. It finds that a simple rule based on minimum teacher distortion consistently selects the same eight‑bit, per‑channel, unclipped configuration, though this does not minimize empirical target cross‑entropy. The study also shows that confidence‑based estimators perform poorly in overconfident regimes, while output‑distribution estimators can outperform the teacher in some architectures, and that combining distortion with a supervised term can improve selection. Across 134 candidate families, teacher‑anchored selection reduces mean regret with very few labels, though the benefit diminishes after about 25 labels.
By Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong
arXiv:2606. 00798v1 Announce Type: cross Abstract: Parameter compression of class-conditional diffusion models reveals an underexplored limitation in output-level distillation: the unconditional score branch remains unsupervised, leaving the classifier-free guidance gap underdetermined in the student.
By Abdullah Al Shafi, Kazi Saeed Alam, Sk Imran Hossain, Engelbert Mephu Nguifo
arXiv:2609.38011v1 Announce Type: new
Abstract: Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downst...
By Diyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco Mondelli
arXiv:2603. 05691v3 Announce Type: replace Abstract: It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models.
By Diyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco Mondelli
arXiv:2510. 14074v2 Announce Type: replace-cross Abstract: We develop a framework for analyzing the learning dynamics of high-dimensional problems trained using one-pass stochastic gradient descent (SGD) with data from multiple anisotropic classes.
By Elizabeth Collins-Woodfin, Inbar Seroussi
arXiv:2607. 14947v1 Announce Type: cross Abstract: Modern generative models are increasingly trained using model-generated signals, creating both opportunities for self-improvement and risks of collapse.
By Saptarshi Roy, Debepsita Mukherjee, Pratik Patil
The paper investigates how the sampled-token reverse-KL loss in on‑policy distillation distributes updates across tokens. By analyzing the gradient of the per‑token K2 estimator, the authors find that tokens with low student probability and large teacher‑student gaps receive disproportionately large gradient norms. They propose Surprise‑aware Reweighting (SuRe), a lightweight weighting rule that further amplifies this allocation, and demonstrate that SuRe improves math metrics on Qwen3 student models without harming out‑of‑domain performance.
By Bing Shao, Jiazheng Zhang, Long Ma, Yujiong Shen, Senjie Jin, Xin Guo, Yuming Yang, Mingxu Chai, Zhiheng Xi, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv:2606. 12171v1 Announce Type: cross Abstract: Knowledge Distillation (KD) and mixup have proven effective at inducing smoothness in class boundaries; KD captures inherent class relationships in probability distributions, and mixup enforces them through convex combinations of inputs.
By Jos\'e Medina, Paul Honeine, Abdelaziz Bensrhair, Amnir Hadachi
arXiv:2609.38342v1 Announce Type: new
Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning....
By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu
arXiv:2606. 25927v1 Announce Type: cross Abstract: As machine learning models and datasets continue to grow, developing complex models has become increasingly computationally demanding.
By Luyang Fang, Haoran Lu, Yongkai Chen, Wenxuan Zhong, Ping Ma