Canalization Before Generalization: Grokking as a Dynamical Probe
Read the original on arXiv Machine Learning →The paper investigates how overparameterized neural networks can fit training data in many different ways yet generalize differently. By applying short, fixed‑duration weight‑decay pulses during the grokking plateau, the authors show that early in training the pulses produce unordered shifts in generalization time, but later a clear dose ordering emerges: larger weight‑decay increases lead to earlier generalization, while larger decreases delay it. This ordering appears before visible generalization and persists even as test‑loss barriers collapse, a phenomenon the authors term the canalization of function selection.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.