Hugging Face Trending Papers

The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression

Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are independent and that their covariance matrices are non-degenerate.

arXiv Machine Learning
Jun 10

Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization

arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.

By Jingfeng Wu, Peter L. Bartlett, Sham M. Kakade, Jason D. Lee, Bin Yu
arXiv Machine Learning
Aug 19

Online Generalized Sparse Regression: How Does Overparametrization Help?

The paper introduces an online generalized-sparsity-constrained regression framework that addresses key challenges in online sparse regression, such as dynamic regularization, memory usage, and real-time computation. It proposes an efficient online hard‑thresholding algorithm that performs closed‑form updates using only summary statistics, achieving global convergence at optimal statistical rates when the projection set is overparameterized. Numerical experiments show the method consistently outperforms existing alternatives in online cardinality‑constrained linear regression and low‑rank matrix sensing.

By Shuoguang Yang, Qiang Sun
arXiv Machine Learning
Aug 19

Causal Discovery in Equal Variance Linear Gaussian DAGs via SURE-Tuned Ridge Regression

The paper introduces SURE-Ridge, a closed‑form estimator for recovering the directed acyclic graph of an equal‑variance linear Gaussian structural equation model. It performs parallel node‑wise regressions with regularization parameters selected via Stein's unbiased risk estimate and then applies adaptive thresholding to produce a DAG from a soft adjacency matrix. Experiments show that SURE‑Ridge attains the lowest structural Hamming distance in small‑sample settings and the fastest run time across all tested sample sizes compared to NOTEARS, DAGMA, and GBNSL.

By Sambit Mishra, Urbashi Mitra
arXiv Machine Learning
Jul 23

Kernel Ridge Regression Inference

arXiv:2302. 06578v4 Announce Type: replace-cross Abstract: We provide uniform confidence bands for kernel ridge regression (KRR), a widely used nonparametric regression estimator for nonstandard data such as preferences, sequences, and graphs.

By Rahul Singh, Suhas Vijaykumar