arXiv Machine Learning By Gengze Xu, Wei Yao, Ziqiao Wang, Yong Liu

Weak-to-Strong Generalization via Bregman Bias-Variance Decomposition

Read the original on arXiv Machine Learning →

The paper studies weak-to-strong generalization (W2SG), where a student model trained on a weaker teacher’s labels surpasses the teacher on the target task. Using a Bregman divergence bias‑variance decomposition, it shows that the student‑teacher risk gap depends on their expected misfit, without requiring convexity of the student hypothesis class. For squared loss, a sufficient condition is that the student converges to the teacher’s posterior mean, achievable by enlarging the student; for cross‑entropy loss, reducing the student’s predictive entropy and using reverse cross‑entropy can promote W2SG, which is empirically validated.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 18

Generalized Kullback-Leibler Divergence Loss

arXiv:2503. 08038v2 Announce Type: replace-cross Abstract: In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss that consists of (1) a weighted Mean Square Error (wMSE) loss and (2) a Cross-Entropy loss incorporating soft labels.

By Jiequan Cui, Beier Zhu, Qingshan Xu, Zhuotao Tian, Xiaojuan Qi, Bei Yu, Hanwang Zhang, Richang Hong
arXiv AI
6d ago

Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift

The paper investigates how to choose the best quantized model from a family of compressed versions when target labels are scarce or unavailable. It finds that a simple rule based on minimum teacher distortion consistently selects the same eight‑bit, per‑channel, unclipped configuration, though this does not minimize empirical target cross‑entropy. The study also shows that confidence‑based estimators perform poorly in overconfident regimes, while output‑distribution estimators can outperform the teacher in some architectures, and that combining distortion with a supervised term can improve selection. Across 134 candidate families, teacher‑anchored selection reduces mean regret with very few labels, though the benefit diminishes after about 25 labels.

By Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong
arXiv AI
Jun 9

Performative Learning Theory

arXiv:2602. 04402v3 Announce Type: replace-cross Abstract: Performative predictions influence the very outcomes they aim to forecast.

By Julian Rodemann, Unai Fischer-Abaigar, James Bailie, Krikamol Muandet