A Smoothed Discrepancy Principle for Random Feature Methods and Neural Networks
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.
arXiv:2406. 04425v2 Announce Type: replace Abstract: A fundamental problem in machine learning is understanding the effect of early stopping on the parameters obtained and the generalization capabilities of the model.
The paper introduces a neighboring early‑stopping rule for adaptive regularization in kernel ridge regression with random features (KRR‑RF). By using a uniform grid in inverse regularization and comparing only adjacent estimators, the method reduces discrepancy checks and can be computed directly in the random‑feature space without forming the full kernel Gram matrix. Under standard source and capacity assumptions, the selected estimator achieves the oracle polynomial learning rate up to logarithmic factors, enabling regularization selection without prior knowledge of smoothness or capacity exponents.
The paper proposes an analytic method for determining the optimal early‑stopping time in training neural networks, avoiding the need for gradient‑descent training. It uses Rademacher complexity with an L1‑norm to estimate generalization error, offering a more general approach than previous random‑matrix‑theory based methods. The framework is demonstrated on linear regression and extended to nonlinear neural networks via linear probing, as shown in a MNIST classification example.
Training neural networks requires balancing the trade-off between fitting the training data and achieving robust performance on unseen inputs. This ability, commonly referred to as generalizability, i...
arXiv:2606. 06772v1 Announce Type: cross Abstract: Understanding the generalization performance of over-parameterized neural networks has become a central topic in deep learning theory.