arXiv Statistics ML

Algebraic Invariants of Lightning Self-Attention

The paper investigates the polynomial coefficients of lightning self‑attention, treating them as coordinates of an algebraic variety. In the single‑token case it identifies the coefficient variety as a rank‑constrained Chow‑type variety and derives algebraic equations; for multiple tokens it shows that linear relations reduce the geometry to coefficients involving interactions between distinct tokens, characterized by a common linear factor and a low‑rank condition. The authors provide explicit families of determinantal, Veronese‑type, and Sylvester resultant‑based invariants, and in the rank‑one case give pencil and flattening equations that define the variety set‑theoretically, with small‑dimension computations confirming the theoretical generators.

arXiv Machine Learning
Jun 16

Constraining the outputs of ReLU neural networks

arXiv:2508. 03867v2 Announce Type: replace-cross Abstract: We introduce a class of algebraic varieties naturally associated with ReLU neural networks, arising from the piecewise linear structure of their outputs across activation regions in input space, and the piecewise multilinear structure in parameter space.

By Yulia Alexandr, Guido Mont\'ufar
arXiv Machine Learning
Aug 28

Linear Independence of Polynomial Compositions and Identifiability of Deep Neural Networks

The paper proposes a conjecture that composing a fixed number of distinct nonconstant polynomials with a generic high‑degree polynomial produces linearly independent polynomials, extending Newman–Slater’s theorem. The authors prove the conjecture for two polynomials and for any number when the degrees are bounded, and they show how these results explain the parameter symmetries of deep fully connected neural networks with generic polynomial activations. In particular, for architectures with layer‑specific activations of increasing degree, the conjecture’s proven cases fully characterize the parameter sets that yield the same end‑to‑end network function, and it also resolves the identifiability of shallow polynomial networks.

By Kathl\'en Kohn, Giovanni Luca Marchetti, Alex Massarenti, Massimiliano Mella
arXiv AI
Jun 12

Real-rootedness of the Poincar\'e polynomials of $\overline{\mathcal M}_{0,n}$: an AI-assisted proof

arXiv:2605. 29151v2 Announce Type: replace-cross Abstract: We prove real-rootedness for the Poincar\'e polynomial \[ P_n(t)=\sum_{i=0}^{n-3} \dim H^{2i}(\overline{\mathcal M}_{0,n};\mathbb{Q})t^i \] of the Deligne--Mumford moduli space $\overline{\mathcal M}_{0,n}$ of stable $n$-pointed rational curves, proving a conjecture of Aluffi--Chen--Marcolli.

By Gergely B\'erczi, Young-Hoon Kiem
arXiv Machine Learning
Aug 19

Tight Bounds for Data-driven Multiple Hyper-parameter Tuning with Structured Loss Function

The paper establishes tight pseudo-dimension bounds for data-driven multiple hyper‑parameter tuning with structured loss functions. By refining upper bounds through real algebraic geometry and analyzing invariant connected sign cells, the authors avoid over‑counting and achieve sharper sample complexities. A multi‑regime lower‑bound framework demonstrates that these upper bounds are tight, and the approach is extended to general bi‑level validation‑loss tuning and broader semi‑algebraic applications.

By Anh Tuan Nguyen, Viet Anh Nguyen
arXiv Machine Learning
Jul 16

Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations

arXiv:2607. 13749v1 Announce Type: new Abstract: Neural networks trained on modular arithmetic exhibit grokking, a delayed transition from memorisation to generalisation known to depend on model capacity: too little and the network memorises slowly or not at all, too much and it generalises almost immediately.

By Chon-Fai Kam, Xavier Cadet, Miloud Bessafi, Frederic Cadet
arXiv Machine Learning
Sep 7

The Geometry of Polynomial Group Convolutional Neural Networks

The paper introduces a new mathematical framework for polynomial group convolutional neural networks (PGCNNs) using graded group algebras. It presents two natural parametrizations of the architecture—based on Hadamard and Kronecker products—that are related by a linear map. The authors compute the dimension of the resulting neuromanifold, show it depends only on the number of layers and group size, and describe the general fiber of the Kronecker parametrization, conjecturing a similar description for the Hadamard case, supported by explicit computations for small groups and shallow networks.

By Yacoub Hendi, Daniel Persson, Magdalena Larfors