arXiv Machine Learning
Jul 30

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.

By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv Machine Learning
Aug 27

Adaptive Regularization for Random Features: A Neighboring Early-Stopping Rule with Oracle-Rate Guarantees

The paper introduces a neighboring early‑stopping rule for adaptive regularization in kernel ridge regression with random features (KRR‑RF). By using a uniform grid in inverse regularization and comparing only adjacent estimators, the method reduces discrepancy checks and can be computed directly in the random‑feature space without forming the full kernel Gram matrix. Under standard source and capacity assumptions, the selected estimator achieves the oracle polynomial learning rate up to logarithmic factors, enabling regularization selection without prior knowledge of smoothness or capacity exponents.

By Caixing Wang, Zhibo Chen, Yue Wang
arXiv Machine Learning
Aug 26

A Data-dependent Early Stopping Rule using Rademacher Complexity with L1-norm

The paper proposes an analytic method for determining the optimal early‑stopping time in training neural networks, avoiding the need for gradient‑descent training. It uses Rademacher complexity with an L1‑norm to estimate generalization error, offering a more general approach than previous random‑matrix‑theory based methods. The framework is demonstrated on linear regression and extended to nonlinear neural networks via linear probing, as shown in a MNIST classification example.

By Duy Hoang, Bastien Berret, Olivier Bruneau, Laurent Fribourg