The Advantages of Fresh Sketching for Ridge Regression
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper introduces an online sketched Newton method that uses a generalized accelerated sketch-and-project solver (GAS) to approximate Newton directions efficiently. GAS incorporates Nesterov momentum and a flexible projection metric, achieving accelerated convergence and reduced computational cost. The authors prove asymptotic normality and a functional central limit theorem for the averaged iterates, enabling an online inference procedure via random scaling that yields a pivotal test statistic with a parameter‑free limiting distribution.
arXiv:2508. 21022v3 Announce Type: replace Abstract: Subsampled natural gradient descent (SNG) has been used to enable high-precision scientific machine learning, but standard analyses based on stochastic preconditioning fail to provide insight into realistic small-sample settings.
arXiv:2609. 08136v1 Announce Type: new Abstract: This paper introduces rlaopt, a PyTorch-based package for large-scale optimization and scientific computing using randomized numerical linear algebra (RandNLA).
arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.
The paper introduces a pooling‑ridge estimation method for functional linear regression that handles data observed at discrete times, ranging from sparse to dense designs. By combining pooling strategies with RKHS‑based techniques, the authors achieve minimax‑optimal prediction risk for both scalar‑on‑function and function‑on‑function models. The study identifies distinct phase transitions in convergence behavior, with up to three transitions for function‑on‑function regression, and validates the approach through simulations and real data examples.
The paper introduces a novel technique called "persistence of memory" to enhance stochastic subspace methods for large‑scale optimisation. By using a weakly correlated guidance vector that is refreshed only at wide intervals, the method provides a structured direction for random subspace descent. The authors demonstrate that this guidance can be efficiently computed in sparse or minibatch settings and present the first theoretical analysis of classical SSD methods for sparse functions, showing alignment with low‑lying Hessian eigenvectors near the optimum.