arXiv AI

At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics

arXiv:2607. 06639v1 Announce Type: cross Abstract: On modular arithmetic, a network's embedding keeps compressing for tens of thousands of steps after it has already generalized.

arXiv Machine Learning
Sep 11

Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking

The study investigates the delayed transition from memorization to generalization—known as grokking—in two‑hidden‑layer MLPs trained on modular arithmetic. By exploring 384 hyperparameter configurations, the authors derive a power‑law scaling relation for the onset time of generalization, showing that data complexity dominates over model capacity. A clear phase boundary at weight decay around 1.0 separates grokking from non‑grokking regimes, and weight norm trajectories indicate implicit regularization during the transition.

By Anish Kataria
arXiv AI
2d ago

Misalignment of Low-Loss Regions Causes Grokking

arXiv:2610.00620v1 Announce Type: cross Abstract: Grokking refers to the delayed emergence of validation-set generalization after a model has already overfit the training set. Although first observed...

By Yongding Tian, Zaid Al-Ars, Maksim Kitsak, Peter Hofstee
Hugging Face Trending Papers
Jun 11

Circuit Synchronization Precedes Generalization: Causal Evidence from Fourier Structure in Grokking Transformers

Grokking -- where a transformer on modular arithmetic suddenly transitions from near-chance to near-perfect validation accuracy -- is attributed to a Fourier circuit, but its timing, causal structure, and controllability remain poorly understood. We introduce the Frequency Synchronization Degree (FSD), a normalised, permutation-tested metric for Fourier circuit synchronisation requiring no prior circuit knowledge.

arXiv Machine Learning
5d ago

The Residual Stream's Effective Depth

The paper introduces effective depth (Deff), a scalar diagnostic that treats a transformer’s layer‑wise residual stream as a discrete‑time process and measures how representation similarity decays with layer distance. Across sixteen decoder‑only language models, Deff reveals that most models exhibit a lower similarity decay than the closed‑form reference, indicating correlated residual updates rather than unused depth. The study also shows that this effect is robust to various controls and persists early in training, suggesting Deff is a global accumulated‑state diagnostic rather than a capability score.

By Barak Gahtan, Ido Galil, Alex M. Bronstein