The paper introduces an online sketched Newton method that uses a generalized accelerated sketch-and-project solver (GAS) to approximate Newton directions efficiently. GAS incorporates Nesterov momentum and a flexible projection metric, achieving accelerated convergence and reduced computational cost. The authors prove asymptotic normality and a functional central limit theorem for the averaged iterates, enabling an online inference procedure via random scaling that yields a pivotal test statistic with a parameter‑free limiting distribution.
By Xinchen Du, Elizaveta Rebrova, Micha{\l} Derezi\'{n}ski, Sen Na
arXiv:2508. 21022v3 Announce Type: replace Abstract: Subsampled natural gradient descent (SNG) has been used to enable high-precision scientific machine learning, but standard analyses based on stochastic preconditioning fail to provide insight into realistic small-sample settings.
By Gil Goldshlager, Jiang Hu, Lin Lin
arXiv:2609. 08136v1 Announce Type: new Abstract: This paper introduces rlaopt, a PyTorch-based package for large-scale optimization and scientific computing using randomized numerical linear algebra (RandNLA).
By Pratik Rathore, Zachary Frangella, Parth Nobel, Xuning Hu, Madeleine Udell
arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.
By Jingfeng Wu, Peter L. Bartlett, Sham M. Kakade, Jason D. Lee, Bin Yu
The paper introduces a pooling‑ridge estimation method for functional linear regression that handles data observed at discrete times, ranging from sparse to dense designs. By combining pooling strategies with RKHS‑based techniques, the authors achieve minimax‑optimal prediction risk for both scalar‑on‑function and function‑on‑function models. The study identifies distinct phase transitions in convergence behavior, with up to three transitions for function‑on‑function regression, and validates the approach through simulations and real data examples.
By Shunxing Yan, Fang Yao
The paper introduces a novel technique called "persistence of memory" to enhance stochastic subspace methods for large‑scale optimisation. By using a weakly correlated guidance vector that is refreshed only at wide intervals, the method provides a structured direction for random subspace descent. The authors demonstrate that this guidance can be efficiently computed in sparse or minibatch settings and present the first theoretical analysis of classical SSD methods for sparse functions, showing alignment with low‑lying Hessian eigenvectors near the optimum.
By Subhroshekhar Ghosh, Clement Z. Q. Ng, Pierre-Louis Poirion, Akiko Takeda
arXiv:2606. 23867v1 Announce Type: new Abstract: The exact computation of the Normalized Maximum Likelihood (NML) codelength for regular non-smooth estimators (e.
By Trenton Lau, Gary P. T. Choi
arXiv:2608. 02539v1 Announce Type: cross Abstract: We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator.
By Jos\'e Luis Montiel Olea, Ryan Strong, Amilcar Velez, Zhuoheng Xu, Haomin Yu
arXiv:2609. 13040v1 Announce Type: new Abstract: We study loss-based filtering for finite-sum optimization with a subset of corrupted component functions whose gradients may be highly unreliable.
By Jamie Haddock, Anna Ma, Elizaveta Rebrova
arXiv:2607. 00252v1 Announce Type: new Abstract: We present an algorithm for the group distributionally robust (GDR) least squares problem.
By Naren Sarayu Manoj, Kumar Kshitij Patel
The paper introduces SURE-Ridge, a closed‑form estimator for recovering the directed acyclic graph of an equal‑variance linear Gaussian structural equation model. It performs parallel node‑wise regressions with regularization parameters selected via Stein's unbiased risk estimate and then applies adaptive thresholding to produce a DAG from a soft adjacency matrix. Experiments show that SURE‑Ridge attains the lowest structural Hamming distance in small‑sample settings and the fastest run time across all tested sample sizes compared to NOTEARS, DAGMA, and GBNSL.
By Sambit Mishra, Urbashi Mitra
arXiv:2606. 15832v1 Announce Type: new Abstract: Empirical risk minimization on massive datasets naturally exhibits a nested double finite-sum structure, where $N=nm$ total samples are logically or physically partitioned into $n$ blocks of size $m$ (e.
By Igor Sokolov, Laurent Condat, Peter Richt\'arik