The paper establishes high‑probability bounds on mixed input derivatives for wide random neural networks whose activation derivatives grow factorially, with a focus on anh networks initialized with Xavier weights. For scalar‑output anh networks with Gaussian weights, the authors prove that when the hidden width exceeds a depth‑dependent threshold, the derivative of any order satisfies a bound that is independent of depth for first‑order derivatives and grows at most polynomially with depth for higher‑order mixed derivatives. These results yield high‑probability estimates for the Euclidean Lipschitz constant and weighted Sobolev norms, linking the regularity of network realizations to quasi‑Monte Carlo integration and its potential impact on QMC‑based training.
By Josef Dick, Michael Feischl, Fabian Zehetgruber
arXiv:2608.23877v1 Announce Type: new
Abstract: We prove a depth hierarchy for ReLU neural networks in which every additional ReLU layer can save exponentially many neurons. For every $\ell\geq 3$, a...
By Itay Safran
arXiv:2607. 07778v1 Announce Type: new Abstract: Bubeck, Li and Nagaraj conjectured that, for generic data, any two-layer neural network with $m$ neurons that fits $n$ noisy labels must have Lipschitz constant at least of order $\sqrt{n/m}$, with no restriction on the size of the weights.
By Yitzchak Shmalo
This paper investigates the ρ^p-Lipschitz constants of deep ReLU neural networks with random weights drawn from a He‑style initialization. For zero‑bias networks, it provides high‑probability upper and lower bounds that differ by at most a logarithmic factor in depth, and shows a sharp contrast between the regimes p∈[1,2) and p∈[2,∞], with the former behaving like the Euclidean norm of a Gaussian vector and the latter like its dual norm. The analysis is extended to networks with non‑zero biases from symmetric distributions, yielding bounds that differ by a logarithmic factor in width and a linear factor in depth.
By Sjoerd Dirksen, Patrick Finke, Paul Geuchen, Dominik St\"oger, Felix Voigtlaender
arXiv:2412. 05109v2 Announce Type: replace Abstract: We derive universal approximation results for the class of (countably) $m$-rectifiable measures.
By Erwin Riegler, Alex B\"uhler, Yang Pan, Helmut B\"olcskei
arXiv:2602. 10949v2 Announce Type: replace-cross Abstract: Effective initialization in deep networks requires an understanding of random neural networks.
By Constantin Kogler, Tassilo Schwarz, Samuel Kittle
arXiv:2609.05572v1 Announce Type: new
Abstract: We prove that every strictly positive probability distribution on \(\{-1,1\}^n\) is represented exactly by a sigmoid belief network with finite paramet...
By Gleb Smirnov
The paper investigates how to reduce computation in neural networks by combining one‑shot magnitude pruning in a static setting with early exit in an adaptive setting. In a simplified single‑neuron model it proves a concentration theorem for pruning and introduces a conditional perceptron whose excess error decreases as a power of the compute gap, with the exponent increasing as partial and full computations align. The authors extend these results to deep networks, showing how pruning distortions accumulate with depth and deriving a compute‑accuracy trade‑off for frozen‑backbone early exit under a Gaussian process framework, with numerical simulations supporting the theoretical scaling laws.
By Erdem Koyuncu
arXiv:2609. 03626v1 Announce Type: cross Abstract: Rigorous results show that feedforward neural networks can overcome the curse of dimensionality in the numerical approximation of high-dimensional partial differential equations (PDEs), but comparatively little is known about residual neural networks (ResNets) in the nonlinear PDE setting.
By Ilkhom Mukhammadiev, Diyora Salimova
arXiv:2506. 08764v3 Announce Type: replace Abstract: Deep neural networks are known to suffer from exploding or vanishing gradients as depth increases, a phenomenon closely tied to the spectral behavior of the input-output Jacobian.
By Benjamin Dadoun, Soufiane Hayou, Hanan Salam, Mohamed El Amine Seddik, Pierre Youssef
The paper demonstrates that for a broad class of spiking neuron models, including the leaky integrate‑and‑fire with subtractive reset, any approximation bound proven for multi‑spike networks can be translated to an equivalent single‑spike network with only a linear change in neuron count, and vice versa. This establishes that single‑spike and multi‑spike neural networks possess identical approximation capabilities for general machine learning tasks. Consequently, existing approximation results for single‑spike networks automatically extend to the multi‑spike case.
By Dominik Dold, Philipp Christian Petersen
arXiv:2210. 16286v2 Announce Type: replace Abstract: To understand the training dynamics of neural networks, prior studies have considered the mean-field limit of two-layer neural networks as the width tends to infinity, establishing theoretical guarantees for its convergence under gradient flow training as well as approximation and generalization capabilities.
By Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna