The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression
arXiv:2607. 24041v1 Announce Type: cross Abstract: Over-parameterized linear regression has been widely studied over the last decade.
Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are independent and that their covariance matrices are non-degenerate.
arXiv:2607. 24041v1 Announce Type: cross Abstract: Over-parameterized linear regression has been widely studied over the last decade.
arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.
arXiv:2608. 07281v1 Announce Type: cross Abstract: This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally.
arXiv:2606. 28573v1 Announce Type: new Abstract: Modern machine learning models are trained by optimizing high-dimensional non-convex empirical risk functions.
arXiv:2309. 15769v3 Announce Type: replace-cross Abstract: Recent advances in deep learning have highlighted the phenomenon of benign overfitting in overparameterized statistical models, sparking significant interest in understanding its foundations.
arXiv:2608. 17466v1 Announce Type: cross Abstract: Regularized sparse regression has been extensively studied in the offline setting, but online formulation remains relatively under-explored.
arXiv:2607. 22436v1 Announce Type: cross Abstract: This work addresses the generation of theoretical correlation matrices with prescribed sparsity patterns associated to graph structures.
arXiv:2608. 17132v1 Announce Type: new Abstract: Recovering the directed acyclic graph (DAG) of a structural equation model (SEM) from observational data is a central problem in causal discovery.
arXiv:2608. 02539v1 Announce Type: cross Abstract: We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator.
arXiv:2302. 06578v4 Announce Type: replace-cross Abstract: We provide uniform confidence bands for kernel ridge regression (KRR), a widely used nonparametric regression estimator for nonstandard data such as preferences, sequences, and graphs.
arXiv:2607. 07888v1 Announce Type: new Abstract: This paper studies distributed sketching for ordinary least squares (OLS) regression, an approach that distributes small sketches of a large data set over multiple machines to separately construct OLS estimators and average them.
arXiv:2607. 08380v1 Announce Type: new Abstract: An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian.