arXiv:2501. 18530v3 Announce Type: replace-cross Abstract: We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width $k$ and input dimension $d$ are large and proportional.
By Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
arXiv:2505. 11766v4 Announce Type: replace Abstract: Neural Operators (NOs) are powerful architectures for learning mappings between function spaces.
By Haoze Song, Zhihao Li, Xiaobo Zhang, Zecheng Gan, Zhilu Lai, Wei Wang
arXiv:2607. 23397v1 Announce Type: new Abstract: Hierarchical neural networks are widely used in artificial intelligence, yet their mathematical properties remain incompletely understood.
By Sumio Watanabe
arXiv:2609.15355v1 Announce Type: cross
Abstract: We study the uniform approximation of smooth scalar-valued functionals on an infinite-dimensional separable Hilbert space by deep ReLU neural network...
By Shuhao Jiao
arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.
By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv:2606. 14954v1 Announce Type: cross Abstract: We develop a general framework for analyzing representation costs of parametric data-fitting methods through their parameter-space regularizers.
By Greg Ongie, Rahul Parhi
arXiv:2608. 09558v1 Announce Type: new Abstract: How expressive is prompting a transformer?
By Alexander Hsu, Rongjie Lai
arXiv:2609. 03129v1 Announce Type: cross Abstract: Several classical machine-learning methods, such as KRRs and SVRs, are both computationally and analytically tractable since their estimators either admit closed-form expressions or are obtained by minimizing convex training objectives; neither feature is generally available for deep neural networks.
By Ruiyang Hong, Hrad Ghoukasian, Anastasis Kratsios
The paper studies operator learning on function spaces using encoder–decoder architectures. It shows that as input and output resolutions grow, the induced kernels converge to a limiting kernel, enabling regularity assumptions independent of resolution. The authors derive upper and lower bounds for regularized stochastic gradient descent, extend the analysis to neural networks via the limiting neural tangent kernel, and provide error bounds and complexity guarantees for various kernel and encoding constructions.
By Lei Shi, Jia-Qi Yang, Ding-Xuan Zhou
arXiv:2608.24007v1 Announce Type: new
Abstract: Understanding how neural networks learn and organize features is central to understanding their behavior. Much existing theory of feature learning has...
By Amirhesam Abedsoltan, Enric Boix-Adsera, Fivos Kalogiannis, Mikhail Belkin
arXiv:2607. 23110v1 Announce Type: cross Abstract: In this paper, we study the extension of Neural Jump ODEs to infinite-dimensional function spaces.
By Florian Krach, Oliver L\"othgren, Josef Teichmann
arXiv:2609.09556v1 Announce Type: new
Abstract: Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear access...
By Enrico Vompa