arXiv Machine Learning

Adaptive Regularization for Random Features: A Neighboring Early-Stopping Rule with Oracle-Rate Guarantees

The paper introduces a neighboring early‑stopping rule for adaptive regularization in kernel ridge regression with random features (KRR‑RF). By using a uniform grid in inverse regularization and comparing only adjacent estimators, the method reduces discrepancy checks and can be computed directly in the random‑feature space without forming the full kernel Gram matrix. Under standard source and capacity assumptions, the selected estimator achieves the oracle polynomial learning rate up to logarithmic factors, enabling regularization selection without prior knowledge of smoothness or capacity exponents.

arXiv Machine Learning
Aug 27

Loop Corrections in Random Feature Models: Training Error and Generalization Gap

The paper investigates fixed‑design random feature ridge regression beyond the mean‑kernel approximation, focusing on how the predictor’s nonlinear dependence on the empirical kernel affects training error, test error, and the conditional generalization gap. By deriving covariance‑level (one‑loop) corrections via a finite resolvent identity, the authors avoid an almost‑sure Neumann‑series assumption and provide explicit remainder bounds for training error. Numerical experiments demonstrate that including mixed train–test covariance tensors is essential for accurate test error predictions and reveal a width–regularization boundary where second‑order truncation becomes unreliable.

By Taeyoung Kim
arXiv AI
Jun 26

XMSE-Aware Adaptive Empirical Bayes Estimation

arXiv:2606. 26975v1 Announce Type: cross Abstract: Empirical Bayes (EB) estimators can match the first-order asymptotic risk of maximum likelihood (ML) while behaving very differently at second order: recent excess mean squared error (XMSE) analysis shows that kernel-based EB estimation may be worse than ML when the kernel is poorly aligned with the true parameter.

By Minghao Chen, Jiale Zheng
arXiv Machine Learning
Jul 7

Distribution-free Deviation Bounds and The Role of Domain Knowledge in Learning via Model Selection with Cross-validation Risk Estimation

arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.

By Diego Marcondes, Cl\'audia Peixoto
Hugging Face Trending Papers
Jun 25

XMSE-Aware Adaptive Empirical Bayes Estimation

Empirical Bayes (EB) estimators can match the first-order asymptotic risk of maximum likelihood (ML) while behaving very differently at second order: recent excess mean squared error (XMSE) analysis shows that kernel-based EB estimation may be worse than ML when the kernel is poorly aligned with the true parameter. This paper turns that diagnostic into a design principle.

arXiv AI
Sep 4

Spectral Convergence of Random Feature Method in Multiple Dimensions

The paper proves spectral convergence of the random feature method (RFM) for multidimensional targets across various function classes, providing high‑probability approximation estimates that hold simultaneously for all admissible error norms. It extends these results to strong‑ and weak‑form RFM discretizations, yielding convergence guarantees for multidimensional second‑order elliptic boundary value and eigenvalue problems. Additionally, it demonstrates super‑exponential singular‑value decay for Fourier features and exponential decay for tanh features, while establishing corresponding condition‑number lower bounds, highlighting a trade‑off between accuracy and ill‑conditioning.

By Pingbing Ming, Hao Yu
arXiv Machine Learning
Jun 10

Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization

arXiv:2509. 17251v2 Announce Type: replace-cross Abstract: Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems.

By Jingfeng Wu, Peter L. Bartlett, Sham M. Kakade, Jason D. Lee, Bin Yu
arXiv Machine Learning
Aug 21

Exact Algebraic Computation of Learning Coefficients for Two-Dimensional Singular Models

arXiv:2608. 20183v1 Announce Type: new Abstract: Classical information criteria such as the Bayesian Information Criterion (BIC) rely on regularity assumptions that break down for singular models, leading to incorrect model selection in settings such as deep learning.

By Gr\'egoire Sergeant-Perthuis (CQSB, Sorbonne Universit\'e), Elias Tsigaridas (Ouragan Team, INRIA), Jules Tsukahara (Ouragan Team, INRIA)