arXiv Machine Learning

The Advantages of Fresh Sketching for Ridge Regression

arXiv Machine Learning
Sep 14

Inference for Newton Methods with Accelerated Sketch-and-Project via Random Scaling

The paper introduces an online sketched Newton method that uses a generalized accelerated sketch-and-project solver (GAS) to approximate Newton directions efficiently. GAS incorporates Nesterov momentum and a flexible projection metric, achieving accelerated convergence and reduced computational cost. The authors prove asymptotic normality and a functional central limit theorem for the averaged iterates, enabling an online inference procedure via random scaling that yields a pivotal test statistic with a parameter‑free limiting distribution.

By Xinchen Du, Elizaveta Rebrova, Micha{\l} Derezi\'{n}ski, Sen Na
arXiv Machine Learning
Jun 10

Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization

arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.

By Jingfeng Wu, Peter L. Bartlett, Sham M. Kakade, Jason D. Lee, Bin Yu
arXiv Machine Learning
Aug 27

Functional linear regression from sparse to dense designs: a pooling-ridge method and minimax optimality

The paper introduces a pooling‑ridge estimation method for functional linear regression that handles data observed at discrete times, ranging from sparse to dense designs. By combining pooling strategies with RKHS‑based techniques, the authors achieve minimax‑optimal prediction risk for both scalar‑on‑function and function‑on‑function models. The study identifies distinct phase transitions in convergence behavior, with up to three transitions for function‑on‑function regression, and validates the approach through simulations and real data examples.

By Shunxing Yan, Fang Yao
arXiv Machine Learning
Sep 17

Gradient Descent with Stochastic Subspaces via Persistence of Memory

The paper introduces a novel technique called "persistence of memory" to enhance stochastic subspace methods for large‑scale optimisation. By using a weakly correlated guidance vector that is refreshed only at wide intervals, the method provides a structured direction for random subspace descent. The authors demonstrate that this guidance can be efficiently computed in sparse or minibatch settings and present the first theoretical analysis of classical SSD methods for sparse functions, showing alignment with low‑lying Hessian eigenvectors near the optimum.

By Subhroshekhar Ghosh, Clement Z. Q. Ng, Pierre-Louis Poirion, Akiko Takeda
arXiv Machine Learning
Aug 19

Causal Discovery in Equal Variance Linear Gaussian DAGs via SURE-Tuned Ridge Regression

The paper introduces SURE-Ridge, a closed‑form estimator for recovering the directed acyclic graph of an equal‑variance linear Gaussian structural equation model. It performs parallel node‑wise regressions with regularization parameters selected via Stein's unbiased risk estimate and then applies adaptive thresholding to produce a DAG from a soft adjacency matrix. Experiments show that SURE‑Ridge attains the lowest structural Hamming distance in small‑sample settings and the fastest run time across all tested sample sizes compared to NOTEARS, DAGMA, and GBNSL.

By Sambit Mishra, Urbashi Mitra