Algebraic Expressivity Certificates for Shallow Polynomial Neural Networks
arXiv:2609. 28500v1 Announce Type: cross Abstract: We study exact representability by bias-free shallow polynomial neural networks using algebraic geometry.
The paper investigates the polynomial coefficients of lightning self‑attention, treating them as coordinates of an algebraic variety. In the single‑token case it identifies the coefficient variety as a rank‑constrained Chow‑type variety and derives algebraic equations; for multiple tokens it shows that linear relations reduce the geometry to coefficients involving interactions between distinct tokens, characterized by a common linear factor and a low‑rank condition. The authors provide explicit families of determinantal, Veronese‑type, and Sylvester resultant‑based invariants, and in the rank‑one case give pencil and flattening equations that define the variety set‑theoretically, with small‑dimension computations confirming the theoretical generators.
arXiv:2609. 28500v1 Announce Type: cross Abstract: We study exact representability by bias-free shallow polynomial neural networks using algebraic geometry.
arXiv:2508. 03867v2 Announce Type: replace-cross Abstract: We introduce a class of algebraic varieties naturally associated with ReLU neural networks, arising from the piecewise linear structure of their outputs across activation regions in input space, and the piecewise multilinear structure in parameter space.
arXiv:2604. 08485v2 Announce Type: replace-cross Abstract: The purpose of this paper is two-fold.
arXiv:2511. 19703v2 Announce Type: replace-cross Abstract: We study the dimension and identifiability of neurovarieties associated to polynomial neural networks.
arXiv:2608.30417v1 Announce Type: new Abstract: We give a complete characterization of equivariant multi-head self-attention (MHSA): if an MHSA layer is equivariant to a symmetry group $G$, then $G$...
arXiv:2609.25776v1 Announce Type: cross Abstract: Nonlinear activations can create equivariant interactions between irreducible representations that linear maps cannot. We use the Gaussian degree dec...
The paper proposes a conjecture that composing a fixed number of distinct nonconstant polynomials with a generic high‑degree polynomial produces linearly independent polynomials, extending Newman–Slater’s theorem. The authors prove the conjecture for two polynomials and for any number when the degrees are bounded, and they show how these results explain the parameter symmetries of deep fully connected neural networks with generic polynomial activations. In particular, for architectures with layer‑specific activations of increasing degree, the conjecture’s proven cases fully characterize the parameter sets that yield the same end‑to‑end network function, and it also resolves the identifiability of shallow polynomial networks.
arXiv:2607. 18817v1 Announce Type: cross Abstract: Algebraic statistics characterizes statistical models through polynomial constraints, but it has mainly been used for analytically specified model classes.
arXiv:2605. 29151v2 Announce Type: replace-cross Abstract: We prove real-rootedness for the Poincar\'e polynomial \[ P_n(t)=\sum_{i=0}^{n-3} \dim H^{2i}(\overline{\mathcal M}_{0,n};\mathbb{Q})t^i \] of the Deligne--Mumford moduli space $\overline{\mathcal M}_{0,n}$ of stable $n$-pointed rational curves, proving a conjecture of Aluffi--Chen--Marcolli.
The paper establishes tight pseudo-dimension bounds for data-driven multiple hyper‑parameter tuning with structured loss functions. By refining upper bounds through real algebraic geometry and analyzing invariant connected sign cells, the authors avoid over‑counting and achieve sharper sample complexities. A multi‑regime lower‑bound framework demonstrates that these upper bounds are tight, and the approach is extended to general bi‑level validation‑loss tuning and broader semi‑algebraic applications.
arXiv:2607. 13749v1 Announce Type: new Abstract: Neural networks trained on modular arithmetic exhibit grokking, a delayed transition from memorisation to generalisation known to depend on model capacity: too little and the network memorises slowly or not at all, too much and it generalises almost immediately.
The paper introduces a new mathematical framework for polynomial group convolutional neural networks (PGCNNs) using graded group algebras. It presents two natural parametrizations of the architecture—based on Hadamard and Kronecker products—that are related by a linear map. The authors compute the dimension of the resulting neuromanifold, show it depends only on the number of layers and group size, and describe the general fiber of the Kronecker parametrization, conjecturing a similar description for the Hadamard case, supported by explicit computations for small groups and shallow networks.