arXiv Machine Learning

The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression

arXiv:2607. 24041v1 Announce Type: cross Abstract: Over-parameterized linear regression has been widely studied over the last decade.

arXiv Machine Learning
Jun 10

Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization

arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.

By Jingfeng Wu, Peter L. Bartlett, Sham M. Kakade, Jason D. Lee, Bin Yu
arXiv Statistics ML
Aug 25

Stochastic gradient descent with initial regularization

The paper studies a variant of stochastic gradient descent called SGDIR, which incorporates initial regularization. It derives dimension‑free upper bounds on the expected excess risk for the squared loss, providing new rates for both averaged and non‑averaged SGDIR under various assumptions. The authors also establish matching lower bounds in certain regimes and compare SGDIR to ridge regression in noisy settings, showing comparable performance up to a polylogarithmic factor.

By Nabil Kahal\'e
arXiv Machine Learning
Jul 14

Exact Dynamics of Multi-class Stochastic Gradient Descent

arXiv:2510. 14074v2 Announce Type: replace-cross Abstract: We develop a framework for analyzing the learning dynamics of high-dimensional problems trained using one-pass stochastic gradient descent (SGD) with data from multiple anisotropic classes.

By Elizabeth Collins-Woodfin, Inbar Seroussi
arXiv Machine Learning
Aug 19

Causal Discovery in Equal Variance Linear Gaussian DAGs via SURE-Tuned Ridge Regression

The paper introduces SURE-Ridge, a closed‑form estimator for recovering the directed acyclic graph of an equal‑variance linear Gaussian structural equation model. It performs parallel node‑wise regressions with regularization parameters selected via Stein's unbiased risk estimate and then applies adaptive thresholding to produce a DAG from a soft adjacency matrix. Experiments show that SURE‑Ridge attains the lowest structural Hamming distance in small‑sample settings and the fastest run time across all tested sample sizes compared to NOTEARS, DAGMA, and GBNSL.

By Sambit Mishra, Urbashi Mitra
arXiv Machine Learning
Jul 7

Efficient Cross-Validation for Sparse Linear Regression

arXiv:2306. 14851v5 Announce Type: replace-cross Abstract: Given a high-dimensional covariate matrix and a response vector, ridge-regularized sparse linear regression selects a subset of features that explains the relationship between covariates and the response in an interpretable manner.

By Ryan Cory-Wright, Andr\'es G\'omez