The paper investigates fixed‑design random feature ridge regression beyond the mean‑kernel approximation, focusing on how the predictor’s nonlinear dependence on the empirical kernel affects training error, test error, and the conditional generalization gap. By deriving covariance‑level (one‑loop) corrections via a finite resolvent identity, the authors avoid an almost‑sure Neumann‑series assumption and provide explicit remainder bounds for training error. Numerical experiments demonstrate that including mixed train–test covariance tensors is essential for accurate test error predictions and reveal a width–regularization boundary where second‑order truncation becomes unreliable.
By Taeyoung Kim
arXiv:2606. 26975v1 Announce Type: cross Abstract: Empirical Bayes (EB) estimators can match the first-order asymptotic risk of maximum likelihood (ML) while behaving very differently at second order: recent excess mean squared error (XMSE) analysis shows that kernel-based EB estimation may be worse than ML when the kernel is poorly aligned with the true parameter.
By Minghao Chen, Jiale Zheng
arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.
By Diego Marcondes, Cl\'audia Peixoto
arXiv:2607. 07513v1 Announce Type: new Abstract: Self-supervised learning matches supervised accuracy from a fraction of the labels, but the labeled-sample efficiency behind this has lacked a theoretical explanation.
By Adam M. Oberman
Empirical Bayes (EB) estimators can match the first-order asymptotic risk of maximum likelihood (ML) while behaving very differently at second order: recent excess mean squared error (XMSE) analysis shows that kernel-based EB estimation may be worse than ML when the kernel is poorly aligned with the true parameter. This paper turns that diagnostic into a design principle.
arXiv:2406. 04425v2 Announce Type: replace Abstract: A fundamental problem in machine learning is understanding the effect of early stopping on the parameters obtained and the generalization capabilities of the model.
By Rishi Sonthalia, Jackie Lok, Elizaveta Rebrova
The paper proves spectral convergence of the random feature method (RFM) for multidimensional targets across various function classes, providing high‑probability approximation estimates that hold simultaneously for all admissible error norms. It extends these results to strong‑ and weak‑form RFM discretizations, yielding convergence guarantees for multidimensional second‑order elliptic boundary value and eigenvalue problems. Additionally, it demonstrates super‑exponential singular‑value decay for Fourier features and exponential decay for tanh features, while establishing corresponding condition‑number lower bounds, highlighting a trade‑off between accuracy and ill‑conditioning.
By Pingbing Ming, Hao Yu
arXiv:2606. 27298v1 Announce Type: cross Abstract: We study the fundamental problem of learning a high-dimensional Gaussian truncated to an unknown halfspace.
By Haitong Liu, Deepak Narayanan Sridharan, David Steurer, Manuel Wiedmer
arXiv:2604. 25021v2 Announce Type: replace Abstract: We study online regression with the square loss in a reproducing kernel Hilbert space under a dynamic regret criterion.
By Dmitry B. Rokhlin, Georgiy A. Karapetyants
arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.
By Jingfeng Wu, Peter L. Bartlett, Sham M. Kakade, Jason D. Lee, Bin Yu
arXiv:2506. 11336v2 Announce Type: replace Abstract: We study the sample complexity of stochastic convex optimization when problem parameters such as the distance to optimality and the Lipschitz constant are unknown.
By Jared Lawrence, Ari Kalinsky, Hannah Bradfield, Yair Carmon, Oliver Hinder
arXiv:2608. 20183v1 Announce Type: new Abstract: Classical information criteria such as the Bayesian Information Criterion (BIC) rely on regularity assumptions that break down for singular models, leading to incorrect model selection in settings such as deep learning.
By Gr\'egoire Sergeant-Perthuis (CQSB, Sorbonne Universit\'e), Elias Tsigaridas (Ouragan Team, INRIA), Jules Tsukahara (Ouragan Team, INRIA)