The paper investigates how many linear samples are needed to learn Lipschitz operators under Gaussian measures. It establishes both lower and upper bounds on the Hermite polynomial approximation error and shows that the minimal worst‑case error cannot converge algebraically with the number of samples. However, if the covariance operator of the Gaussian measure decays rapidly, convergence rates arbitrarily close to any algebraic rate can be achieved.
By Ben Adcock, Michael Griebel, Gregor Maier
arXiv:2603. 00819v2 Announce Type: replace-cross Abstract: This paper surveys recent developments at the intersection of operator learning, statistical learning theory, and approximation theory.
By Simone Brugiapaglia, Nicola Rares Franco, Nicholas H. Nelsen
arXiv:2606. 01244v2 Announce Type: replace-cross Abstract: Inspired by the function-space theory of neural networks, we formulate and analyze a variation space for nonlinear operators between Hilbert spaces, defined through vector-valued Borel measures of bounded variation.
By Jia-Qi Yang, Lei Shi
The paper studies operator learning on function spaces using encoder–decoder architectures. It shows that as input and output resolutions grow, the induced kernels converge to a limiting kernel, enabling regularity assumptions independent of resolution. The authors derive upper and lower bounds for regularized stochastic gradient descent, extend the analysis to neural networks via the limiting neural tangent kernel, and provide error bounds and complexity guarantees for various kernel and encoding constructions.
By Lei Shi, Jia-Qi Yang, Ding-Xuan Zhou
arXiv:2504.18184v5 Announce Type: replace
Abstract: We consider a class of statistical inverse problems involving the estimation of a regression operator from a Polish space to a separable Hilbert sp...
By Jia-Qi Yang, Lei Shi
arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.
By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv:2402.04691v5 Announce Type: replace-cross
Abstract: This study investigates the use of stochastic gradient descent (SGD) to learn operators between general Hilbert spaces. We study weak and str...
By Lei Shi, Jia-Qi Yang
arXiv:2608. 15982v1 Announce Type: new Abstract: We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces.
By Mahdi Mohammadigohari, Thomas Borsani, Giuseppe Di Fatta
arXiv:2606. 14954v1 Announce Type: cross Abstract: We develop a general framework for analyzing representation costs of parametric data-fitting methods through their parameter-space regularizers.
By Greg Ongie, Rahul Parhi
arXiv:2606. 17419v1 Announce Type: new Abstract: We develop approximation and generalization error estimates for multi-input neural operators, with the output error measured in Sobolev norms.
By Yahong Yang, Zecheng Zhang, Wei Zhu, Wenjing Liao, Hao Liu
The paper investigates how many data samples per domain are needed for effective learning across multiple domains. It derives criteria from learning bounds that reveal an inverse linear relationship between the number of training domains and the required samples per domain, offering theoretical guidance for dataset adequacy and construction. The study also establishes a close link between in-domain learning and out-of-domain generalization through new generalization bounds.
By Hong Zheng
This paper investigates the ρ^p-Lipschitz constants of deep ReLU neural networks with random weights drawn from a He‑style initialization. For zero‑bias networks, it provides high‑probability upper and lower bounds that differ by at most a logarithmic factor in depth, and shows a sharp contrast between the regimes p∈[1,2) and p∈[2,∞], with the former behaving like the Euclidean norm of a Gaussian vector and the latter like its dual norm. The analysis is extended to networks with non‑zero biases from symmetric distributions, yielding bounds that differ by a logarithmic factor in width and a linear factor in depth.
By Sjoerd Dirksen, Patrick Finke, Paul Geuchen, Dominik St\"oger, Felix Voigtlaender