Singular Learning and Occam's Razor in Deep Monomial Networks
arXiv:2606. 28464v1 Announce Type: new Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture.
arXiv:2410. 00722v3 Announce Type: replace Abstract: We study convolutional neural networks with monomial activation functions.
arXiv:2606. 28464v1 Announce Type: new Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture.
arXiv:2508. 03867v2 Announce Type: replace-cross Abstract: We introduce a class of algebraic varieties naturally associated with ReLU neural networks, arising from the piecewise linear structure of their outputs across activation regions in input space, and the piecewise multilinear structure in parameter space.
The paper introduces a new mathematical framework for polynomial group convolutional neural networks (PGCNNs) using graded group algebras. It presents two natural parametrizations of the architecture—based on Hadamard and Kronecker products—that are related by a linear map. The authors compute the dimension of the resulting neuromanifold, show it depends only on the number of layers and group size, and describe the general fiber of the Kronecker parametrization, conjecturing a similar description for the Hadamard case, supported by explicit computations for small groups and shallow networks.
arXiv:2604. 14037v2 Announce Type: replace Abstract: Parameter space is not function space for neural network architectures.
The paper proposes a conjecture that composing a fixed number of distinct nonconstant polynomials with a generic high‑degree polynomial produces linearly independent polynomials, extending Newman–Slater’s theorem. The authors prove the conjecture for two polynomials and for any number when the degrees are bounded, and they show how these results explain the parameter symmetries of deep fully connected neural networks with generic polynomial activations. In particular, for architectures with layer‑specific activations of increasing degree, the conjecture’s proven cases fully characterize the parameter sets that yield the same end‑to‑end network function, and it also resolves the identifiability of shallow polynomial networks.
arXiv:2605. 09609v2 Announce Type: replace Abstract: We provide counterexamples to the unimodal minimal filling architecture conjecture for polynomial neural networks (PNNs) with power activation functions.
arXiv:2609.39078v1 Announce Type: new Abstract: Representations are routinely used across machine learning, psychology, and neuroscience to draw inferences about the computations of biological and ar...
arXiv:2511. 19703v2 Announce Type: replace-cross Abstract: We study the dimension and identifiability of neurovarieties associated to polynomial neural networks.
arXiv:2410. 04907v2 Announce Type: replace-cross Abstract: In this paper we contribute to the frequently studied question of how to decompose a continuous piecewise linear (CPWL) function into a difference of two convex CPWL functions.
The paper introduces a general theoretical framework for fibrations on graphs labeled by a commutative monoid, extending the classic theory of graph fibrations to weighted and algebraically labeled graphs. It also accommodates approximate fibrations and demonstrates how this framework can be used to compress arbitrary neural networks, including CNNs, providing a solid theoretical basis for recent findings on fibration symmetries in geometric deep learning.
arXiv:2606. 04327v1 Announce Type: cross Abstract: We investigate the geometric structure of stationary plateaus that arise in the loss landscape of two-layer neural networks with smooth activation functions.
arXiv:2609.25776v1 Announce Type: cross Abstract: Nonlinear activations can create equivariant interactions between irreducible representations that linear maps cannot. We use the Gaussian degree dec...