arXiv:2607. 11666v1 Announce Type: new Abstract: Grokking is a phenomenon in which neural networks initially memorize training data and only later exhibit strong generalization after prolonged optimization.
By Maksim A Kazanskii
arXiv:2511. 01938v3 Announce Type: replace-cross Abstract: Grokking is a puzzling phenomenon in neural networks where full generalization occurs only after a substantial delay following the complete memorization of the training data.
By Tiberiu Musat
arXiv:2607. 23967v1 Announce Type: new Abstract: Delayed generalization, or grokking, remains poorly understood despite extensive empirical study.
By Taeyoung Kim
The study investigates the delayed transition from memorization to generalization—known as grokking—in two‑hidden‑layer MLPs trained on modular arithmetic. By exploring 384 hyperparameter configurations, the authors derive a power‑law scaling relation for the onset time of generalization, showing that data complexity dominates over model capacity. A clear phase boundary at weight decay around 1.0 separates grokking from non‑grokking regimes, and weight norm trajectories indicate implicit regularization during the transition.
By Anish Kataria
arXiv:2609.07755v1 Announce Type: new
Abstract: Understanding generalization remains a central challenge in machine learning because it requires jointly considering data, architecture, and training d...
By Yuqing Wang, Ioannis G. Kevrekidis, Mikhail Belkin
arXiv:2608. 01833v1 Announce Type: cross Abstract: Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization.
By Lai Shun Chan, Xiaotian Zhang, Yue Shang, Ge Zhang, Entao Yang
arXiv:2606. 30512v1 Announce Type: cross Abstract: Why overparameterised deep networks generalise so remarkably well remains one of the most stubborn open questions in machine learning theory.
By Srinivasa Rao P., Vangmayi P Reddy
The paper derives an exact discrete‑time law that captures how learning‑rate schedules and weight decay interact in scale‑invariant neural networks, showing that a single scalar quantity governs the effective step size. It demonstrates that the balance point between contraction and expansion is intrinsically unstable, leading to recurrent dynamics when using constant learning rates with weight decay. The authors extend this analysis to various optimizers and datasets, confirming the law’s precision and showing that performance peaks sharply at the predicted boundary.
By Hasan Amin, Wei-Kai Chang, Rajiv Khanna
The paper investigates how normalization makes neural networks scale‑invariant, creating a feedback loop between learning‑rate schedules and weight decay that controls the effective step size of the optimizer. It derives an exact discrete‑time law showing that a single scalar quantity captures all schedule and decay effects, with norm growth providing a self‑quenching counter‑force that defines a sharp boundary between contraction‑ and expansion‑dominated regimes. Through exact analysis of a normalized regression model and experiments on MLPs, CNNs, GPT‑2, and various datasets, the authors demonstrate that constant learning rates with weight decay are intrinsically unstable, leading to recurrent dynamics, and that adaptive optimizers exhibit weaker stabilization under normalization.
"whyItMatters":"The study provides a precise, actionable rule for controlling training dynamics and schedule design in modern deep learning by isolating a single governing quantity for scale‑invariant optimization."
The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.
By Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands
arXiv:2606. 13753v1 Announce Type: cross Abstract: Grokking is the delayed onset of generalization in neural networks, arising long after they fit the training data.
By Truong Xuan Khanh, Doan Hoang Viet, Luu Duc Trung, Phan Thanh Duc
arXiv:2608. 14803v1 Announce Type: new Abstract: A recent line of work recasts the post-memorization phase of grokking as constrained optimization: once a network interpolates the training set, weight decay drives a slow drift along the zero-loss manifold toward lower norm.
By Suvinava Basak