arXiv Machine Learning

Variation Spaces for Encoder--Decoder Neural Operators: Approximation and Generalization

arXiv:2606. 01244v2 Announce Type: replace-cross Abstract: Inspired by the function-space theory of neural networks, we formulate and analyze a variation space for nonlinear operators between Hilbert spaces, defined through vector-valued Borel measures of bounded variation.

arXiv Machine Learning
2d ago

Resolution-Independent Analysis of Encoder--Decoder Operator Learning via Limiting Kernels

The paper studies operator learning on function spaces using encoder–decoder architectures. It shows that as input and output resolutions grow, the induced kernels converge to a limiting kernel, enabling regularity assumptions independent of resolution. The authors derive upper and lower bounds for regularized stochastic gradient descent, extend the analysis to neural networks via the limiting neural tangent kernel, and provide error bounds and complexity guarantees for various kernel and encoding constructions.

By Lei Shi, Jia-Qi Yang, Ding-Xuan Zhou
arXiv Machine Learning
Jul 16

New universal operator approximation theorem for encoder-decoder architectures

arXiv:2503. 24092v2 Announce Type: replace-cross Abstract: Motivated by the rapidly growing field of mathematics for operator approximation with neural networks, we present a novel universal operator approximation theorem for broad classes of encoder-decoder architectures and a wide range of input and output spaces.

By Janek G\"odeke, Pascal Fernsel
arXiv Machine Learning
Aug 18

Operator-Theoretic Generalization Bounds for Multitask Deep Learning

arXiv:2608. 15982v1 Announce Type: new Abstract: We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces.

By Mahdi Mohammadigohari, Thomas Borsani, Giuseppe Di Fatta
Hugging Face Trending Papers
Aug 17

Operator-Theoretic Generalization Bounds for Multitask Deep Learning

We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces. In vector-valued Sobolev RKHSs, we derive Rademacher complexity bounds for invertible and width-expanding injective architectures.

arXiv Machine Learning
Sep 7

The Sample Complexity of Learning Lipschitz Operators with respect to Gaussian Measures

The paper investigates how many linear samples are needed to learn Lipschitz operators under Gaussian measures. It establishes both lower and upper bounds on the Hermite polynomial approximation error and shows that the minimal worst‑case error cannot converge algebraically with the number of samples. However, if the covariance operator of the Gaussian measure decays rapidly, convergence rates arbitrarily close to any algebraic rate can be achieved.

By Ben Adcock, Michael Griebel, Gregor Maier
arXiv Machine Learning
Sep 4

A Closed-Form Formula for Consistent Lipschitz Regression on Metric Spaces with Sparse Neural Network Realizations

arXiv:2609. 03129v1 Announce Type: cross Abstract: Several classical machine-learning methods, such as KRRs and SVRs, are both computationally and analytically tractable since their estimators either admit closed-form expressions or are obtained by minimizing convex training objectives; neither feature is generally available for deep neural networks.

By Ruiyang Hong, Hrad Ghoukasian, Anastasis Kratsios
arXiv Machine Learning
6d ago

Near-optimal estimates for the $\ell^p$-Lipschitz constants of deep random ReLU neural networks

This paper investigates the ρ^p-Lipschitz constants of deep ReLU neural networks with random weights drawn from a He‑style initialization. For zero‑bias networks, it provides high‑probability upper and lower bounds that differ by at most a logarithmic factor in depth, and shows a sharp contrast between the regimes p∈[1,2) and p∈[2,∞], with the former behaving like the Euclidean norm of a Gaussian vector and the latter like its dual norm. The analysis is extended to networks with non‑zero biases from symmetric distributions, yielding bounds that differ by a logarithmic factor in width and a linear factor in depth.

By Sjoerd Dirksen, Patrick Finke, Paul Geuchen, Dominik St\"oger, Felix Voigtlaender