arXiv AI

The Geometry of Grokking: Norm Minimization on the Zero-Loss Manifold

arXiv:2511. 01938v3 Announce Type: replace-cross Abstract: Grokking is a puzzling phenomenon in neural networks where full generalization occurs only after a substantial delay following the complete memorization of the training data.

arXiv Machine Learning
1d ago

Learning the identity: a case study of how SGD selects among functional decompositions

The paper investigates how stochastic gradient descent (SGD) selects specific functional decompositions when training a deep linear residual network to learn the identity function. Although many weight configurations minimize the population loss, SGD consistently prefers particular solutions, especially under anisotropic label noise or different parametrizations. The authors explain this bias using an entropic loss term that penalizes the expected squared norm of the minibatch gradient, analytically characterizing its minimizers and showing that trained networks align with these predictions.

By Andy Arditi, Weian Xie, David Bau, Liu Ziyin
arXiv AI
Jun 30

A Stochastic--Geometric Theory of Scaling Laws in Grokking

arXiv:2606. 30388v1 Announce Type: cross Abstract: Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only begins to generalize after a prolonged delay, often through an abrupt transition.

By R\'ois\'in Luo, Christian Gagn\'e, Jonas Ngnaw\'e, Ihsan Ullah, Karyn Morrissey
arXiv Machine Learning
Sep 10

The Dynamics of Generalization in Deep Learning

arXiv:2504.16450v4 Announce Type: replace Abstract: We derive a differential equation that governs the evolution of the generalization gap when a model is trained by gradient descent-based methods. T...

By Rubing Yang, Pratik Chaudhari
arXiv Machine Learning
Sep 23

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.

By Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands