arXiv:2608. 07281v1 Announce Type: cross Abstract: This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally.
By Zhijun Liu, Dandan Jiang
arXiv:2109.02355v2 Announce Type: replace
Abstract: The last decade of progress in machine learning (ML), especially the deep learning era, has raised a number of scientific questions that challenge...
By Yehuda Dar, Vidya Muthukumar, Richard G. Baraniuk
The paper develops a statistical theory for minimum‑norm interpolation in high‑dimensional regression, showing how regularization geometry and signal sparsity affect generalization. It identifies regimes where sparsity‑promoting regularizers yield exact interpolation that is far more accurate than approximate fitting, and proves a zero–one generalization law for strongly overparameterized noiseless problems. The authors also characterize training and generalization errors along ρ‑regularization paths when feature dimension and sample size are proportional, demonstrating that generalization improves with more sparsity‑promoting norms and sparser targets, and that small changes in regularization strength can cause large shifts in generalization.
whyItMatters":"The work provides a quantitative understanding of delayed generalization (grokking) and reveals a statistical instability in minimum‑norm interpolation, offering insights that could guide the design of regularizers for better generalization in overparameterized models."
By Gil Kur, Ileana Rugina, Cl\'ementine Carla Juliette Domin\'e, Marco Mondelli
arXiv:2601. 19791v4 Announce Type: replace Abstract: We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting.
By Mingyue Xu, Gal Vardi, Itay Safran
arXiv:2604. 08625v2 Announce Type: replace-cross Abstract: We develop a theoretical framework for generalization in the interpolating regime of statistical learning.
By Gustav Olaf Yunus Laitinen-Lundstr\"om Fredriksson-Imanov
The paper investigates why diffusion models, unlike typical deep learning models, exhibit catastrophic overfitting when overparameterized. Through experiments on U‑Nets trained on CelebA and a random‑features theoretical analysis, it shows that the interpolation peak occurs at a model size proportional to the product of training samples and noise realizations, but the test loss starts to rise already at the number of samples, leading to memorization of the empirical score. Regularization techniques such as ridge penalties or early stopping can still make large models outperform smaller, unregularized ones.
By Rapha\"el Urfin, Tony Bonnaire, Giulio Biroli, Marc M\'ezard
The paper examines whether benign overfitting—where highly overparameterized models still predict well—occurs in equity return prediction. It finds a double‑descent risk curve for ridgeless models and shows that while ridge regularization slightly improves performance, the advantage vanishes at high parameter‑to‑observation ratios. Ultimately, both models fail to beat a simple historical average, indicating that standard equity predictors lack genuine forecasting power even with flexible machine learning methods.
By Hui Guo, Jiawei Huang, Runze Li, Yan Yu
Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularizatio...
arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.
By Diego Marcondes, Cl\'audia Peixoto
arXiv:2406. 13944v2 Announce Type: replace-cross Abstract: This paper establishes the generalization error of pooled min-$\ell_2$-norm interpolation in transfer learning, where data from diverse distributions are available.
By Yanke Song, Kenneth Gu, Sohom Bhattacharya, Pragya Sur
arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.
By Jingfeng Wu, Peter L. Bartlett, Sham M. Kakade, Jason D. Lee, Bin Yu
arXiv:2604. 13130v2 Announce Type: replace Abstract: We study learning to learn through the lens of hyperparameter tuning.
By Saumya Goyal, Rohith Rongali, Ritabrata Ray, Barnab\'as P\'oczos