arXiv Machine Learning

Stochastic Gradient Descent for Operator Learning in Hilbert Spaces: Convergence Rates and Minimax Lower Bounds

arXiv Machine Learning
Aug 26

Sequential operator learning under dependent data

The paper presents time‑uniform self‑normalized concentration bounds for stochastic processes in Hilbert spaces with vector‑valued noise, enabling regression‑error guarantees for both linear and nonlinear parametric operators. These results apply to possibly infinite‑dimensional inputs and outputs without requiring independence or mixing assumptions, and are derived in the context of sequentially collected, dependent data such as adaptive experimental design and dynamical‑system modelling.

By Rafael Oliveira
arXiv Machine Learning
Jul 30

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.

By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv Machine Learning
3d ago

Neural Operators for Nonlinear Functionals on RKHS

arXiv:2403.12187v2 Announce Type: replace-cross Abstract: Motivated by the abundance of functional data, such as time series and images, we study the approximation and statistical learning of nonline...

By Tian-Yi Zhou, Namjoon Suh, Guang Cheng, Xiaoming Huo
arXiv Machine Learning
Sep 7

The Sample Complexity of Learning Lipschitz Operators with respect to Gaussian Measures

The paper investigates how many linear samples are needed to learn Lipschitz operators under Gaussian measures. It establishes both lower and upper bounds on the Hermite polynomial approximation error and shows that the minimal worst‑case error cannot converge algebraically with the number of samples. However, if the covariance operator of the Gaussian measure decays rapidly, convergence rates arbitrarily close to any algebraic rate can be achieved.

By Ben Adcock, Michael Griebel, Gregor Maier
arXiv Machine Learning
3d ago

Resolution-Independent Analysis of Encoder--Decoder Operator Learning via Limiting Kernels

The paper studies operator learning on function spaces using encoder–decoder architectures. It shows that as input and output resolutions grow, the induced kernels converge to a limiting kernel, enabling regularity assumptions independent of resolution. The authors derive upper and lower bounds for regularized stochastic gradient descent, extend the analysis to neural networks via the limiting neural tangent kernel, and provide error bounds and complexity guarantees for various kernel and encoding constructions.

By Lei Shi, Jia-Qi Yang, Ding-Xuan Zhou
arXiv Machine Learning
Aug 26

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

The paper investigates diffusion models trained in a lazy high‑dimensional regime, extending benign overfitting theory to generative settings. By analyzing denoising score matching in a vector‑valued RKHS with an inner‑product kernel, the authors derive exact risk trajectories under gradient flow when the number of samples scales proportionally with dimensionality. These trajectories reveal three distinct phases—spectral generalization, noise‑dominated interpolation, and empirical Bayes memorization—whose interplay shapes the distribution of generated samples.

By Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz