arXiv:2607. 06639v1 Announce Type: cross Abstract: On modular arithmetic, a network's embedding keeps compressing for tens of thousands of steps after it has already generalized.
By Truong Xuan Khanh
arXiv:2601. 19791v4 Announce Type: replace Abstract: We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting.
By Mingyue Xu, Gal Vardi, Itay Safran
arXiv:2607. 05104v1 Announce Type: cross Abstract: Grokking -- the delayed onset of generalization long after a network has fit its training set - -is usually studied in models too large to read completely and reported from single training runs.
By Yoshiyuki Ootani
The study investigates the delayed transition from memorization to generalization—known as grokking—in two‑hidden‑layer MLPs trained on modular arithmetic. By exploring 384 hyperparameter configurations, the authors derive a power‑law scaling relation for the onset time of generalization, showing that data complexity dominates over model capacity. A clear phase boundary at weight decay around 1.0 separates grokking from non‑grokking regimes, and weight norm trajectories indicate implicit regularization during the transition.
By Anish Kataria
arXiv:2607. 11666v1 Announce Type: new Abstract: Grokking is a phenomenon in which neural networks initially memorize training data and only later exhibit strong generalization after prolonged optimization.
By Maksim A Kazanskii
arXiv:2606. 05863v1 Announce Type: new Abstract: Grokking suggests that fitting the training data and learning a simple underlying rule may occur on different time scales.
By Hu Tan, Kuo Gai, Shihua Zhang
arXiv:2608. 07436v1 Announce Type: new Abstract: Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head.
By Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau, Anass Belfatmi
arXiv:2607. 29503v1 Announce Type: new Abstract: While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is.
By Xiaotian Zhang, Lai Shun Chan, Yue Shang, Entao Yang, Ge Zhang
arXiv:2606. 17120v1 Announce Type: new Abstract: Deep neural networks (DNNs) exhibit first order phase transitions under variations of the L2 regularization strength, with each transition marking the onset of a new learnable feature.
By Ibrahim Talha Ersoy, Karoline Wiesner
The paper investigates why the train‑validation performance gap widens during fine‑tuning of pretrained models. It proposes a dynamic structural explanation: as training proceeds, updates shift from broadly reusable features to more example‑specific ones, increasing gradient heterogeneity and the gap. Experiments on synthetic ResMLP hierarchies, NLP models (RoBERTa, DeBERTa, Qwen) across six datasets, and vision models (ResNet‑18) confirm that higher reliance on private features correlates with larger accuracy gaps, supporting the proposed account.
By Yuchen Li, Mingyu Du, Zongqi Fan, Ken-Tye Yong, Nguyen H. Tran
Grokking -- where a transformer on modular arithmetic suddenly transitions from near-chance to near-perfect validation accuracy -- is attributed to a Fourier circuit, but its timing, causal structure, and controllability remain poorly understood. We introduce the Frequency Synchronization Degree (FSD), a normalised, permutation-tested metric for Fourier circuit synchronisation requiring no prior circuit knowledge.
arXiv:2604. 00316v2 Announce Type: replace-cross Abstract: Grokking occurs when a model achieves high training accuracy but generalization to unseen test points happens long after that.
By Marcel Tom\`as Bernal, Neil Rohit Mallinar, Mikhail Belkin