arXiv:2609.09130v1 Announce Type: new
Abstract: An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input...
By Xiaoyu Li, Zhizhou Sha, Jiaojiao Jiang, Junbin Gao, Andi Han
arXiv:2609.06327v2 Announce Type: replace-cross
Abstract: A query-oblivious coreset for a softmax-attention head is a subset of the key-value pairs whose attention output is within $\varepsilon$ of t...
By Ofek I. Cohen
arXiv:2607. 07778v1 Announce Type: new Abstract: Bubeck, Li and Nagaraj conjectured that, for generic data, any two-layer neural network with $m$ neurons that fits $n$ noisy labels must have Lipschitz constant at least of order $\sqrt{n/m}$, with no restriction on the size of the weights.
By Yitzchak Shmalo
arXiv:2510. 04060v3 Announce Type: replace-cross Abstract: We establish two related but logically distinct results for shallow ReLU$^k$ neural networks on the unit sphere $\SS^d$.
By Tong Mao, Jinchao Xu
This paper investigates the ρ^p-Lipschitz constants of deep ReLU neural networks with random weights drawn from a He‑style initialization. For zero‑bias networks, it provides high‑probability upper and lower bounds that differ by at most a logarithmic factor in depth, and shows a sharp contrast between the regimes p∈[1,2) and p∈[2,∞], with the former behaving like the Euclidean norm of a Gaussian vector and the latter like its dual norm. The analysis is extended to networks with non‑zero biases from symmetric distributions, yielding bounds that differ by a logarithmic factor in width and a linear factor in depth.
By Sjoerd Dirksen, Patrick Finke, Paul Geuchen, Dominik St\"oger, Felix Voigtlaender
The paper establishes high‑probability bounds on mixed input derivatives for wide random neural networks whose activation derivatives grow factorially, with a focus on anh networks initialized with Xavier weights. For scalar‑output anh networks with Gaussian weights, the authors prove that when the hidden width exceeds a depth‑dependent threshold, the derivative of any order satisfies a bound that is independent of depth for first‑order derivatives and grows at most polynomially with depth for higher‑order mixed derivatives. These results yield high‑probability estimates for the Euclidean Lipschitz constant and weighted Sobolev norms, linking the regularity of network realizations to quasi‑Monte Carlo integration and its potential impact on QMC‑based training.
By Josef Dick, Michael Feischl, Fabian Zehetgruber
The paper investigates restricted eigenvalue (RE) bounds for norm‑regularized estimators under heavy‑tailed designs. It shows that the previously conjectured sample‑size law based on Gaussian width fails for heavy‑tailed measurements, due to a phenomenon called simultaneous threshold occupancy. The authors provide explicit counterexamples, derive worst‑case sample‑complexity bounds, and compare the behavior of heavy‑tailed versus Gaussian designs on constant‑width polyhedral descent cones.
By Shi Fu, Huibo Xu, Qixin Zhang, Dacheng Tao
arXiv:2607. 22889v1 Announce Type: new Abstract: Learning the natural parameters $z \in \mathbb{R}^n$ of discrete distributions $\mu_z$ from independent samples constrained to a subset $S \subseteq \{0,1\}^n$ is a foundational challenge in high-dimensional statistics.
By Rohan Chauhan, Ioannis Panageas
Uniform stability controls how much one training example can change the loss at any test point. A new logarithmic-free upper bound shows that a $γ$-uniformly stable algorithm with loss in $[0,L]$ has...
arXiv:2608. 15472v1 Announce Type: cross Abstract: The problem of networked information aggregation, studied in Kearns et al.
By Ambar Pal
A persistent empirical observation is that trained neural networks outperform their neural tangent kernel (NTK) limit on tasks with compositional structure, yet a quantitative account of $\textbf{when}$ and $\textbf{by how much}$ has been lacking. Working on the unit circle, we give such an account through a dichotomy between two complexity measures of the target: its $\textbf{Fourier complexity}$, which controls NTK kernel regression, and its $\textbf{architectural complexity}$, which controls learning over depth-$L$, width-$w$ ReLU networks with the variation norm of the weights bounded by $R$.
arXiv:2607. 06382v1 Announce Type: cross Abstract: A persistent empirical observation is that trained neural networks outperform their neural tangent kernel (NTK) limit on tasks with compositional structure, yet a quantitative account of $\textbf{when}$ and $\textbf{by how much}$ has been lacking.
By Arkaprabha Ganguli, Emil Constantinescu