arXiv Machine Learning

A Smoothed Discrepancy Principle for Random Feature Methods and Neural Networks

arXiv Machine Learning
Jul 30

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.

By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv Machine Learning
Aug 27

Adaptive Regularization for Random Features: A Neighboring Early-Stopping Rule with Oracle-Rate Guarantees

The paper introduces a neighboring early‑stopping rule for adaptive regularization in kernel ridge regression with random features (KRR‑RF). By using a uniform grid in inverse regularization and comparing only adjacent estimators, the method reduces discrepancy checks and can be computed directly in the random‑feature space without forming the full kernel Gram matrix. Under standard source and capacity assumptions, the selected estimator achieves the oracle polynomial learning rate up to logarithmic factors, enabling regularization selection without prior knowledge of smoothness or capacity exponents.

By Caixing Wang, Zhibo Chen, Yue Wang
arXiv Machine Learning
Aug 26

A Data-dependent Early Stopping Rule using Rademacher Complexity with L1-norm

The paper proposes an analytic method for determining the optimal early‑stopping time in training neural networks, avoiding the need for gradient‑descent training. It uses Rademacher complexity with an L1‑norm to estimate generalization error, offering a more general approach than previous random‑matrix‑theory based methods. The framework is demonstrated on linear regression and extended to nonlinear neural networks via linear probing, as shown in a MNIST classification example.

By Duy Hoang, Bastien Berret, Olivier Bruneau, Laurent Fribourg
arXiv Machine Learning
Jul 9

Fixed-Gaussian Spectral Algorithms: Minimax Optimal Rates for Misspecified Learning and Transfer

arXiv:2501. 10870v2 Announce Type: replace-cross Abstract: The principal objective of this work is twofold within nonparametric regression settings: (1) to establish the minimax optimal convergence rates for fixed-bandwidth Gaussian kernel spectral algorithms when the true regression function resides in a Sobolev space, and (2) to apply Gaussian spectral algorithms for achieving robust and adaptive transfer learning under concept shift.

By Haotian Lin, Matthew Reimherr
arXiv AI
Sep 24

Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon

The paper presents a near-complete, nonasymptotic generalization theory for multilayer neural networks using path regularization, applicable to broad Lipschitz loss functions without requiring bounded loss or extreme network hyperparameters. It provides an explicit upper bound that addresses approximation rates in generalized Barron spaces and demonstrates the double descent phenomenon for ReLU networks. The authors claim near-minimax optimality for regression problems and plan to establish matching lower bounds in future work.

By Hao Yu
arXiv Machine Learning
Aug 27

Loop Corrections in Random Feature Models: Training Error and Generalization Gap

The paper investigates fixed‑design random feature ridge regression beyond the mean‑kernel approximation, focusing on how the predictor’s nonlinear dependence on the empirical kernel affects training error, test error, and the conditional generalization gap. By deriving covariance‑level (one‑loop) corrections via a finite resolvent identity, the authors avoid an almost‑sure Neumann‑series assumption and provide explicit remainder bounds for training error. Numerical experiments demonstrate that including mixed train–test covariance tensors is essential for accurate test error predictions and reveal a width–regularization boundary where second‑order truncation becomes unreliable.

By Taeyoung Kim