arXiv Machine Learning

How to Tame Grokking: Representation Geometry as a Control Signal

arXiv:2607. 11666v1 Announce Type: new Abstract: Grokking is a phenomenon in which neural networks initially memorize training data and only later exhibit strong generalization after prolonged optimization.

arXiv AI
Jun 30

A Stochastic--Geometric Theory of Scaling Laws in Grokking

arXiv:2606. 30388v1 Announce Type: cross Abstract: Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only begins to generalize after a prolonged delay, often through an abrupt transition.

By R\'ois\'in Luo, Christian Gagn\'e, Jonas Ngnaw\'e, Ihsan Ullah, Karyn Morrissey
arXiv Machine Learning
Sep 23

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.

By Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands
arXiv Machine Learning
Sep 25

Pointwise Generalization in Deep Neural Networks

The paper introduces a pointwise generalization theory for fully connected deep neural networks, using a pointwise Riemannian Dimension derived from eigenvalues of learned feature representations across layers. This framework provides hypothesis-dependent, representation-aware generalization bounds that are significantly tighter than traditional size- or norm-based approaches, both theoretically and experimentally. The authors analytically identify structural properties that explain deep networks’ tractability and empirically show that the pointwise Riemannian Dimension captures feature compression, over‑parameterization effects, and optimizer bias.

By Shaojie Li, Yunbei Xu
arXiv AI
Jun 17

Dimensionality Controls When Modularity Helps in Continual Learning

arXiv:2606. 17889v1 Announce Type: cross Abstract: Compositional learning systems must balance plasticity, the ability to acquire new knowledge, with stability, the preservation of previously learned components, especially when tasks share structure and risk interference.

By Kathrin Korte, Christian Medeiros Adriano, Joachim Winther Pedersen, Eleni Nisioti, Sebastian Risi