On the Geometry and Optimization of Polynomial Convolutional Networks
arXiv:2410. 00722v3 Announce Type: replace Abstract: We study convolutional neural networks with monomial activation functions.
arXiv:2607. 08370v1 Announce Type: cross Abstract: We derive bounds for the volume of tubular neighbourhoods of smooth Pfaffian hypersurfaces, generalising known results for algebraic varieties.
arXiv:2410. 00722v3 Announce Type: replace Abstract: We study convolutional neural networks with monomial activation functions.
arXiv:2605. 09609v2 Announce Type: replace Abstract: We provide counterexamples to the unimodal minimal filling architecture conjecture for polynomial neural networks (PNNs) with power activation functions.
arXiv:2511. 19703v2 Announce Type: replace-cross Abstract: We study the dimension and identifiability of neurovarieties associated to polynomial neural networks.
arXiv:2603. 12785v2 Announce Type: replace Abstract: Three-layer neural networks are known to form singular learning models, and their Bayesian asymptotic behavior is governed by the learning coefficient, or real log canonical threshold.
The paper proposes a conjecture that composing a fixed number of distinct nonconstant polynomials with a generic high‑degree polynomial produces linearly independent polynomials, extending Newman–Slater’s theorem. The authors prove the conjecture for two polynomials and for any number when the degrees are bounded, and they show how these results explain the parameter symmetries of deep fully connected neural networks with generic polynomial activations. In particular, for architectures with layer‑specific activations of increasing degree, the conjecture’s proven cases fully characterize the parameter sets that yield the same end‑to‑end network function, and it also resolves the identifiability of shallow polynomial networks.
arXiv:2508. 03867v2 Announce Type: replace-cross Abstract: We introduce a class of algebraic varieties naturally associated with ReLU neural networks, arising from the piecewise linear structure of their outputs across activation regions in input space, and the piecewise multilinear structure in parameter space.
arXiv:2606. 28464v1 Announce Type: new Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture.
arXiv:2607. 23397v1 Announce Type: new Abstract: Hierarchical neural networks are widely used in artificial intelligence, yet their mathematical properties remain incompletely understood.
arXiv:2607. 04597v1 Announce Type: new Abstract: In this paper, we study the universal approximation property of residual neural networks, and obtain some new results.
arXiv:2602. 12390v2 Announce Type: replace Abstract: We study neural networks with trainable low-degree rational activation functions and show that they are more expressive and parameter-efficient than modern piecewise-linear and smooth activations such as ELU, LeakyReLU, LogSigmoid, PReLU, ReLU, SELU, CELU, Sigmoid, SiLU, Mish, Softplus, Tanh, Softmin, Softmax, and LogSoftmax.
arXiv:2607. 06781v1 Announce Type: new Abstract: In this work, we investigate the fixed-architecture neural network approximation with explicit parameter bounds and elementary activations.
arXiv:2609. 19937v1 Announce Type: cross Abstract: Recent studies have shown that smooth functions can be well approximated by ReLU neural networks with path norm constraint on the weights.