arXiv Machine Learning

Loop Corrections in Random Feature Models: Training Error and Generalization Gap

The paper investigates fixed‑design random feature ridge regression beyond the mean‑kernel approximation, focusing on how the predictor’s nonlinear dependence on the empirical kernel affects training error, test error, and the conditional generalization gap. By deriving covariance‑level (one‑loop) corrections via a finite resolvent identity, the authors avoid an almost‑sure Neumann‑series assumption and provide explicit remainder bounds for training error. Numerical experiments demonstrate that including mixed train–test covariance tensors is essential for accurate test error predictions and reveal a width–regularization boundary where second‑order truncation becomes unreliable.

arXiv Machine Learning
Aug 27

Adaptive Regularization for Random Features: A Neighboring Early-Stopping Rule with Oracle-Rate Guarantees

The paper introduces a neighboring early‑stopping rule for adaptive regularization in kernel ridge regression with random features (KRR‑RF). By using a uniform grid in inverse regularization and comparing only adjacent estimators, the method reduces discrepancy checks and can be computed directly in the random‑feature space without forming the full kernel Gram matrix. Under standard source and capacity assumptions, the selected estimator achieves the oracle polynomial learning rate up to logarithmic factors, enabling regularization selection without prior knowledge of smoothness or capacity exponents.

By Caixing Wang, Zhibo Chen, Yue Wang
arXiv Machine Learning
Jun 9

Generalization in Nonlinear Least Squares via Learned Feature Geometry

arXiv:2606. 08799v1 Announce Type: cross Abstract: We study the generalization of ridge-regularized nonlinear least-squares models via on-average algorithmic stability, deriving error bounds for local minimizers in terms of a data-dependent effective dimension that reflects the geometry of the gradient model at the trained parameters, through the empirical Jacobian Gram matrix and a residual--curvature term.

By Ayub Kharel, Ilja Kuzborski, Patrick Rebeschini, Yasin Abbasi-Yadkori
arXiv Machine Learning
4d ago

Structured Features Overfit Where Random Features Grok

arXiv:2609. 15047v1 Announce Type: new Abstract: Xu, Vardi and Safran (ICML 2026) prove that over-parameterized ridge regression over an unstructured random Gaussian feature map groks, with the delay between memorization and generalization growing as $1/\lambda$ in the weight decay.

By Chon-Fai Kam, Miloud Bessafi, Frederic Cadet