arXiv AI By Tiberiu Musat

The Geometry of Grokking: Norm Minimization on the Zero-Loss Manifold

Read the original on arXiv AI →

arXiv:2511. 01938v3 Announce Type: replace-cross Abstract: Grokking is a puzzling phenomenon in neural networks where full generalization occurs only after a substantial delay following the complete memorization of the training data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

Learning the identity: a case study of how SGD selects among functional decompositions

The paper investigates how stochastic gradient descent (SGD) selects specific functional decompositions when training a deep linear residual network to learn the identity function. Although many weight configurations minimize the population loss, SGD consistently prefers particular solutions, especially under anisotropic label noise or different parametrizations. The authors explain this bias using an entropic loss term that penalizes the expected squared norm of the minibatch gradient, analytically characterizing its minimizers and showing that trained networks align with these predictions.

By Andy Arditi, Weian Xie, David Bau, Liu Ziyin
arXiv AI
Jun 30

A Stochastic--Geometric Theory of Scaling Laws in Grokking

arXiv:2606. 30388v1 Announce Type: cross Abstract: Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only begins to generalize after a prolonged delay, often through an abrupt transition.

By R\'ois\'in Luo, Christian Gagn\'e, Jonas Ngnaw\'e, Ihsan Ullah, Karyn Morrissey