arXiv AI By Truong Xuan Khanh

What Does the Weight Norm Control in Grokking? Logit-Scale Mediation under Cross-Entropy

Read the original on arXiv AI →

arXiv:2606. 18465v1 Announce Type: cross Abstract: Grokking, the delayed jump from memorization to generalization, is usually tied to the weight norm: a smaller norm generalizes sooner.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 11

Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking

The study investigates the delayed transition from memorization to generalization—known as grokking—in two‑hidden‑layer MLPs trained on modular arithmetic. By exploring 384 hyperparameter configurations, the authors derive a power‑law scaling relation for the onset time of generalization, showing that data complexity dominates over model capacity. A clear phase boundary at weight decay around 1.0 separates grokking from non‑grokking regimes, and weight norm trajectories indicate implicit regularization during the transition.

By Anish Kataria