arXiv:2609.07755v1 Announce Type: new
Abstract: Understanding generalization remains a central challenge in machine learning because it requires jointly considering data, architecture, and training d...
By Yuqing Wang, Ioannis G. Kevrekidis, Mikhail Belkin
arXiv:2505. 22578v2 Announce Type: replace Abstract: The optimization of neural networks under weight decay remains poorly understood from a theoretical standpoint.
By Etienne Boursier, Matthew Bowditch, Matthias Englert, Ranko Lazic
The paper investigates how stochastic gradient descent (SGD) selects specific functional decompositions when training a deep linear residual network to learn the identity function. Although many weight configurations minimize the population loss, SGD consistently prefers particular solutions, especially under anisotropic label noise or different parametrizations. The authors explain this bias using an entropic loss term that penalizes the expected squared norm of the minibatch gradient, analytically characterizing its minimizers and showing that trained networks align with these predictions.
By Andy Arditi, Weian Xie, David Bau, Liu Ziyin
arXiv:2606. 05863v1 Announce Type: new Abstract: Grokking suggests that fitting the training data and learning a simple underlying rule may occur on different time scales.
By Hu Tan, Kuo Gai, Shihua Zhang
arXiv:2606. 30388v1 Announce Type: cross Abstract: Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only begins to generalize after a prolonged delay, often through an abrupt transition.
By R\'ois\'in Luo, Christian Gagn\'e, Jonas Ngnaw\'e, Ihsan Ullah, Karyn Morrissey
arXiv:2601. 19791v4 Announce Type: replace Abstract: We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting.
By Mingyue Xu, Gal Vardi, Itay Safran