Theory-to-Practice Gap for Neural Networks and Neural Operators
arXiv:2503. 18219v2 Announce Type: replace Abstract: This work studies the sampling complexity of learning with ReLU neural networks and neural operators.
arXiv:2603. 00819v2 Announce Type: replace-cross Abstract: This paper surveys recent developments at the intersection of operator learning, statistical learning theory, and approximation theory.
arXiv:2503. 18219v2 Announce Type: replace Abstract: This work studies the sampling complexity of learning with ReLU neural networks and neural operators.
arXiv:2402.04691v5 Announce Type: replace-cross Abstract: This study investigates the use of stochastic gradient descent (SGD) to learn operators between general Hilbert spaces. We study weak and str...
arXiv:2609.36709v1 Announce Type: cross Abstract: Out-of-distribution (OOD) generalization is a central challenge in scientific machine learning. We study regression problems in which the test distri...
arXiv:2504.18184v5 Announce Type: replace Abstract: We consider a class of statistical inverse problems involving the estimation of a regression operator from a Polish space to a separable Hilbert sp...
The paper investigates how many linear samples are needed to learn Lipschitz operators under Gaussian measures. It establishes both lower and upper bounds on the Hermite polynomial approximation error and shows that the minimal worst‑case error cannot converge algebraically with the number of samples. However, if the covariance operator of the Gaussian measure decays rapidly, convergence rates arbitrarily close to any algebraic rate can be achieved.
arXiv:2609.39512v1 Announce Type: new Abstract: The small-sample learning problem remains a fundamental challenge in machine learning because limited training data lead to unstable model estimation a...
arXiv:2603. 20388v2 Announce Type: replace-cross Abstract: We derive the asymptotic risk function of regularized empirical risk minimization (ERM) estimators tuned by $n$-fold cross-validation (CV).
The paper presents a near-complete, nonasymptotic generalization theory for multilayer neural networks using path regularization, applicable to broad Lipschitz loss functions without requiring bounded loss or extreme network hyperparameters. It provides an explicit upper bound that addresses approximation rates in generalized Barron spaces and demonstrates the double descent phenomenon for ReLU networks. The authors claim near-minimax optimality for regression problems and plan to establish matching lower bounds in future work.
The paper presents time‑uniform self‑normalized concentration bounds for stochastic processes in Hilbert spaces with vector‑valued noise, enabling regression‑error guarantees for both linear and nonlinear parametric operators. These results apply to possibly infinite‑dimensional inputs and outputs without requiring independence or mixing assumptions, and are derived in the context of sequentially collected, dependent data such as adaptive experimental design and dynamical‑system modelling.
arXiv:2606. 17419v1 Announce Type: new Abstract: We develop approximation and generalization error estimates for multi-input neural operators, with the output error measured in Sobolev norms.
arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.
arXiv:2406. 12264v5 Announce Type: replace-cross Abstract: We obtain a new universal approximation theorem for continuous (possibly nonlinear) operators on arbitrary Banach spaces using the Leray-Schauder mapping.