Canalization Before Generalization: Grokking as a Dynamical Probe
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper investigates how overparameterized neural networks can fit training data in many different ways yet generalize differently. By applying short, fixed‑duration weight‑decay pulses during the grokking plateau, the authors show that early in training the pulses produce unordered shifts in generalization time, but later a clear dose ordering emerges: larger weight‑decay increases lead to earlier generalization, while larger decreases delay it. This ordering appears before visible generalization and persists even as test‑loss barriers collapse, a phenomenon the authors term the canalization of function selection.
arXiv:2607. 23967v1 Announce Type: new Abstract: Delayed generalization, or grokking, remains poorly understood despite extensive empirical study.
arXiv:2601. 19791v4 Announce Type: replace Abstract: We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting.
arXiv:2609.07755v1 Announce Type: new Abstract: Understanding generalization remains a central challenge in machine learning because it requires jointly considering data, architecture, and training d...
arXiv:2607. 20552v1 Announce Type: new Abstract: Grokking -- the delayed generalization of neural networks long after they have memorized their training data -- wastes thousands of training epochs and is notoriously unpredictable.
arXiv:2602.01718v2 Announce Type: replace Abstract: Predicting generalization from quantities available before target-test evaluation remains a central challenge in deep learning. The systematic benc...