arXiv:2607. 07035v1 Announce Type: cross Abstract: The architecture of deep feedforward neural networks is ubiquitous in deep learning, either as a whole system or as a subnetwork of other architectures, and thus its mechanism is a key ingredient of the black box of neural networks.
By Changcun Huang
arXiv:2607. 23397v1 Announce Type: new Abstract: Hierarchical neural networks are widely used in artificial intelligence, yet their mathematical properties remain incompletely understood.
By Sumio Watanabe
arXiv:2402. 00152v5 Announce Type: replace Abstract: Constructing the architecture of a neural network is a challenging pursuit for the machine learning community, and the dilemma of whether to go deeper or wider remains a persistent question.
By Yahong Yang, Juncai He
We develop a convergent scheme to train neural networks involving analytic activation functions based on gradient flows. Convergence properties are guaranteed by Lojasiewicz theory.
arXiv:2607. 13574v1 Announce Type: cross Abstract: We develop a convergent scheme to train neural networks involving analytic activation functions based on gradient flows.
By Ana Carpio
arXiv:2607. 10869v1 Announce Type: new Abstract: We study the population gradient flow of an infinitely wide two-layer neural network learning a misspecified single-index model in high dimension.
By C\'edric Gerbelot, Jean-Christophe Mourrat
arXiv:2608. 14733v1 Announce Type: cross Abstract: Building on the foundation of single-hidden-layer neural networks, Fourier Feature Networks (FENs) are proposed, which incorporate Fourier features using $\cos$, $\sin$, or a combination of both.
By Qihong Yang, Zhijie Su, Yangtao Deng, Qiaolin He
arXiv:2606. 09820v1 Announce Type: cross Abstract: We generalize the universal approximation theorem for functional input neural networks (FNN) to differentiable maps by including the approximation of the derivatives.
By Philipp Schmocker, Josef Teichmann
arXiv:2601. 07397v2 Announce Type: replace-cross Abstract: In this work, we propose a novel layerwise adaptive construction method for neural network architectures.
By Michael Hinterm\"uller, Michael Hinze, Denis Korolev
arXiv:2607. 10200v1 Announce Type: new Abstract: The Neural Tangent Kernel (NTK) is one powerful tool for analyzing the training dynamics of neural networks in the over-parameterized regime.
By Bangti Jin, Longjun Wu
arXiv:2605. 01702v2 Announce Type: replace Abstract: Theoretical studies show that for any differentiable function on a compact domain, there exists a neural network that approximates both the function values and gradients.
By Sejun Park, Yeachan Park, Geonho Hwang
arXiv:2606. 20325v1 Announce Type: new Abstract: Classical approximation theorems ask for a new neural network whenever the target accuracy is improved.
By Valentin Abadie, Clemens Hutter, Helmut B\"olcskei