The paper develops a statistical theory for minimum‑norm interpolation in high‑dimensional regression, showing how regularization geometry and signal sparsity affect generalization. It identifies regimes where sparsity‑promoting regularizers yield exact interpolation that is far more accurate than approximate fitting, and proves a zero–one generalization law for strongly overparameterized noiseless problems. The authors also characterize training and generalization errors along ρ‑regularization paths when feature dimension and sample size are proportional, demonstrating that generalization improves with more sparsity‑promoting norms and sparser targets, and that small changes in regularization strength can cause large shifts in generalization.
whyItMatters":"The work provides a quantitative understanding of delayed generalization (grokking) and reveals a statistical instability in minimum‑norm interpolation, offering insights that could guide the design of regularizers for better generalization in overparameterized models."
By Gil Kur, Ileana Rugina, Cl\'ementine Carla Juliette Domin\'e, Marco Mondelli
The paper investigates why diffusion models, unlike typical deep learning models, exhibit catastrophic overfitting when overparameterized. Through experiments on U‑Nets trained on CelebA and a random‑features theoretical analysis, it shows that the interpolation peak occurs at a model size proportional to the product of training samples and noise realizations, but the test loss starts to rise already at the number of samples, leading to memorization of the empirical score. Regularization techniques such as ridge penalties or early stopping can still make large models outperform smaller, unregularized ones.
By Rapha\"el Urfin, Tony Bonnaire, Giulio Biroli, Marc M\'ezard
Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularizatio...
arXiv:2609.16788v1 Announce Type: new
Abstract: Noise2Noise (N2N) trains denoisers on pairs of independently corrupted observations, eliminating clean references. We stress-test two natural conjectur...
By Dingyan Shang, Zhenyu Xu, Youting Wang, Bonan Shen, Bowen Liu
arXiv:2606. 28573v1 Announce Type: new Abstract: Modern machine learning models are trained by optimizing high-dimensional non-convex empirical risk functions.
By Andrea Montanari, Kangjie Zhou
arXiv:2608. 03197v1 Announce Type: new Abstract: Sharpness-Aware Minimization (SAM) improves generalization by seeking parameters whose loss is robust to local adversarial perturbations, but the quantitative mechanism underlying its implicit bias toward flat minima remains unclear.
By Jiaxin Deng, Junbiao Pang