arXiv:2608. 07281v1 Announce Type: cross Abstract: This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally.
By Zhijun Liu, Dandan Jiang
arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.
By Jingfeng Wu, Peter L. Bartlett, Sham M. Kakade, Jason D. Lee, Bin Yu
The paper studies a variant of stochastic gradient descent called SGDIR, which incorporates initial regularization. It derives dimension‑free upper bounds on the expected excess risk for the squared loss, providing new rates for both averaged and non‑averaged SGDIR under various assumptions. The authors also establish matching lower bounds in certain regimes and compare SGDIR to ridge regression in noisy settings, showing comparable performance up to a polylogarithmic factor.
By Nabil Kahal\'e
arXiv:2607. 01895v1 Announce Type: new Abstract: We study ridge-regularized log-density-ratio estimation in the Gaussian location model with a common covariance matrix.
By Francis Bach (SIERRA)
arXiv:2309. 15769v3 Announce Type: replace-cross Abstract: Recent advances in deep learning have highlighted the phenomenon of benign overfitting in overparameterized statistical models, sparking significant interest in understanding its foundations.
By Dennis Shen, Dogyoon Song, Peng Ding, Jasjeet S. Sekhon
arXiv:2510. 12249v2 Announce Type: replace Abstract: In performative learning, the data distribution reacts to the deployed model - for example, because strategic users adapt their features to game it - which creates a more complex dynamic than in classical supervised learning.
By Edwige Cyffers, Alireza Mirrokni, Marco Mondelli
arXiv:2609.38011v1 Announce Type: new
Abstract: Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downst...
By Diyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco Mondelli
arXiv:2406. 04425v2 Announce Type: replace Abstract: A fundamental problem in machine learning is understanding the effect of early stopping on the parameters obtained and the generalization capabilities of the model.
By Rishi Sonthalia, Jackie Lok, Elizaveta Rebrova
arXiv:2608. 04860v1 Announce Type: cross Abstract: This paper develops procedures for nonparametric goodness-of-fit testing under covariate shift, where labelled data are drawn from a source population but goodness-of-fit is evaluated for a target population.
By Zhen Hou, Dong Xia
arXiv:2603. 05691v3 Announce Type: replace Abstract: It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models.
By Diyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco Mondelli
arXiv:2606. 00322v1 Announce Type: new Abstract: We introduce a perturbative approach for nonparametric instrumental variable (NPIV) estimation.
By Wei Bu, Arthur Gretton
The paper examines whether benign overfitting—where highly overparameterized models still predict well—occurs in equity return prediction. It finds a double‑descent risk curve for ridgeless models and shows that while ridge regularization slightly improves performance, the advantage vanishes at high parameter‑to‑observation ratios. Ultimately, both models fail to beat a simple historical average, indicating that standard equity predictors lack genuine forecasting power even with flexible machine learning methods.
By Hui Guo, Jiawei Huang, Runze Li, Yan Yu