The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression
arXiv:2607. 24041v1 Announce Type: cross Abstract: Over-parameterized linear regression has been widely studied over the last decade.
Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are independent and that their covariance matrices are non-degenerate.
arXiv:2607. 24041v1 Announce Type: cross Abstract: Over-parameterized linear regression has been widely studied over the last decade.
arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.
arXiv:2608. 07281v1 Announce Type: cross Abstract: This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally.
arXiv:2609.39440v1 Announce Type: new Abstract: We compare the instance-wise, finite-sample risks of monotone spectral filters for linear regression, a broad class of estimators including principal c...
arXiv:2606. 28573v1 Announce Type: new Abstract: Modern machine learning models are trained by optimizing high-dimensional non-convex empirical risk functions.
arXiv:2309. 15769v3 Announce Type: replace-cross Abstract: Recent advances in deep learning have highlighted the phenomenon of benign overfitting in overparameterized statistical models, sparking significant interest in understanding its foundations.
The paper introduces an online generalized-sparsity-constrained regression framework that addresses key challenges in online sparse regression, such as dynamic regularization, memory usage, and real-time computation. It proposes an efficient online hard‑thresholding algorithm that performs closed‑form updates using only summary statistics, achieving global convergence at optimal statistical rates when the projection set is overparameterized. Numerical experiments show the method consistently outperforms existing alternatives in online cardinality‑constrained linear regression and low‑rank matrix sensing.
arXiv:2607. 22436v1 Announce Type: cross Abstract: This work addresses the generation of theoretical correlation matrices with prescribed sparsity patterns associated to graph structures.
The paper introduces SURE-Ridge, a closed‑form estimator for recovering the directed acyclic graph of an equal‑variance linear Gaussian structural equation model. It performs parallel node‑wise regressions with regularization parameters selected via Stein's unbiased risk estimate and then applies adaptive thresholding to produce a DAG from a soft adjacency matrix. Experiments show that SURE‑Ridge attains the lowest structural Hamming distance in small‑sample settings and the fastest run time across all tested sample sizes compared to NOTEARS, DAGMA, and GBNSL.
arXiv:2608. 02539v1 Announce Type: cross Abstract: We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator.
Recovering the directed acyclic graph (DAG) of a structural equation model (SEM) from observational data is a central problem in causal discovery. The iterative gradient descent and per-problem hyperp...
arXiv:2302. 06578v4 Announce Type: replace-cross Abstract: We provide uniform confidence bands for kernel ridge regression (KRR), a widely used nonparametric regression estimator for nonstandard data such as preferences, sequences, and graphs.