Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are independent and that their covariance matrices are non-degenerate.
arXiv:2608. 07281v1 Announce Type: cross Abstract: This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally.
By Zhijun Liu, Dandan Jiang
arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.
By Jingfeng Wu, Peter L. Bartlett, Sham M. Kakade, Jason D. Lee, Bin Yu
arXiv:2609.39440v1 Announce Type: new
Abstract: We compare the instance-wise, finite-sample risks of monotone spectral filters for linear regression, a broad class of estimators including principal c...
By Juno Kim, Hengyu Fu, Peter Bartlett, Jason D. Lee, Jingfeng Wu
arXiv:2309. 15769v3 Announce Type: replace-cross Abstract: Recent advances in deep learning have highlighted the phenomenon of benign overfitting in overparameterized statistical models, sparking significant interest in understanding its foundations.
By Dennis Shen, Dogyoon Song, Peng Ding, Jasjeet S. Sekhon
arXiv:2609.38565v1 Announce Type: new
Abstract: Over the past 25 years, sketching and sampling have become widely used tools for accelerating large-scale regression. In iterative randomized solvers,...
By Linkai Ma, Qilin Li, Petros Drineas
arXiv:2606. 28573v1 Announce Type: new Abstract: Modern machine learning models are trained by optimizing high-dimensional non-convex empirical risk functions.
By Andrea Montanari, Kangjie Zhou
The paper studies a variant of stochastic gradient descent called SGDIR, which incorporates initial regularization. It derives dimension‑free upper bounds on the expected excess risk for the squared loss, providing new rates for both averaged and non‑averaged SGDIR under various assumptions. The authors also establish matching lower bounds in certain regimes and compare SGDIR to ridge regression in noisy settings, showing comparable performance up to a polylogarithmic factor.
By Nabil Kahal\'e
arXiv:2510. 14074v2 Announce Type: replace-cross Abstract: We develop a framework for analyzing the learning dynamics of high-dimensional problems trained using one-pass stochastic gradient descent (SGD) with data from multiple anisotropic classes.
By Elizabeth Collins-Woodfin, Inbar Seroussi
The paper introduces SURE-Ridge, a closed‑form estimator for recovering the directed acyclic graph of an equal‑variance linear Gaussian structural equation model. It performs parallel node‑wise regressions with regularization parameters selected via Stein's unbiased risk estimate and then applies adaptive thresholding to produce a DAG from a soft adjacency matrix. Experiments show that SURE‑Ridge attains the lowest structural Hamming distance in small‑sample settings and the fastest run time across all tested sample sizes compared to NOTEARS, DAGMA, and GBNSL.
By Sambit Mishra, Urbashi Mitra
arXiv:2608. 02539v1 Announce Type: cross Abstract: We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator.
By Jos\'e Luis Montiel Olea, Ryan Strong, Amilcar Velez, Zhuoheng Xu, Haomin Yu
arXiv:2306. 14851v5 Announce Type: replace-cross Abstract: Given a high-dimensional covariate matrix and a response vector, ridge-regularized sparse linear regression selects a subset of features that explains the relationship between covariates and the response in an interpretable manner.
By Ryan Cory-Wright, Andr\'es G\'omez