arXiv Machine Learning By Gunn Kim

Dynamical phase selection controls compute scaling in looped transformers

Read the original on arXiv Machine Learning →

The paper investigates looped transformers, which perform inference by repeatedly applying a weight‑tied map, making their computation a dynamical process. It shows that identical architectures trained to the same accuracy can converge to distinct dynamical phases—one governed by a saddle‑node fold and another by a Neimark‑Sacker transition—each with different compute scaling behaviors. The study derives a parameter‑free relation linking relaxation time and spectral gap in the fold phase and demonstrates how critical slowing down leads to a workload‑level tail distribution, while the Neimark‑Sacker phase eliminates this scaling law.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 23

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.

By Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands
arXiv Machine Learning
Sep 11

Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking

The study investigates the delayed transition from memorization to generalization—known as grokking—in two‑hidden‑layer MLPs trained on modular arithmetic. By exploring 384 hyperparameter configurations, the authors derive a power‑law scaling relation for the onset time of generalization, showing that data complexity dominates over model capacity. A clear phase boundary at weight decay around 1.0 separates grokking from non‑grokking regimes, and weight norm trajectories indicate implicit regularization during the transition.

By Anish Kataria