arXiv Machine Learning By Luca Muscarnera, Silas Ruhrberg Est\'evez, Yuanzhang Xiao, Mihaela Van der Schaar

Fast Generalization after Interpolation via Critically Damped Momentum Optimization

Read the original on arXiv Machine Learning →

arXiv:2606. 01521v1 Announce Type: new Abstract: A central problem in machine learning is that models can achieve near-perfect training performance while generalizing substantially less well to unseen examples.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Statistics ML
3d ago

Grokking through the Lens of Minimum-Norm Interpolation

The paper develops a statistical theory for minimum‑norm interpolation in high‑dimensional regression, showing how regularization geometry and signal sparsity affect generalization. It identifies regimes where sparsity‑promoting regularizers yield exact interpolation that is far more accurate than approximate fitting, and proves a zero–one generalization law for strongly overparameterized noiseless problems. The authors also characterize training and generalization errors along ρ‑regularization paths when feature dimension and sample size are proportional, demonstrating that generalization improves with more sparsity‑promoting norms and sparser targets, and that small changes in regularization strength can cause large shifts in generalization. whyItMatters":"The work provides a quantitative understanding of delayed generalization (grokking) and reveals a statistical instability in minimum‑norm interpolation, offering insights that could guide the design of regularizers for better generalization in overparameterized models."

By Gil Kur, Ileana Rugina, Cl\'ementine Carla Juliette Domin\'e, Marco Mondelli