On Universality of Deep Equivariant Networks
arXiv:2510. 15814v2 Announce Type: replace-cross Abstract: Universality results for equivariant neural networks remain rare.
arXiv:2406. 08966v3 Announce Type: replace Abstract: The separation power of a machine learning model refers to its ability to distinguish between different inputs and is often used as a proxy for its expressivity.
arXiv:2510. 15814v2 Announce Type: replace-cross Abstract: Universality results for equivariant neural networks remain rare.
arXiv:2606. 02490v1 Announce Type: new Abstract: This work studies neural architectures for classifying symmetric positive-definite matrices, focusing on congruence-like layers, in which the input matrix is multiplied on the left and right by a (possibly rectangular) weight matrix $W$ and its transpose.
This work studies neural architectures for classifying symmetric positive-definite matrices, focusing on congruence-like layers, in which the input matrix is multiplied on the left and right by a (possibly rectangular) weight matrix $W$ and its transpose. Such layers lie at the core of the celebrated SPDNet and have also been employed independently for dimensionality reduction on positive-definite data.
The paper introduces the directional linear separability measure (D‑LSM) to quantify how much linear separability is preserved or improved by injective affine maps in neural networks. It characterizes the geometry supporting D‑LSM, proves its invariance under injective affine embeddings, and derives conditions for gated activations (ReLU, GELU, SiLU) to preserve and recover samples. Experiments validate the theoretical bounds, demonstrate affine‑tube constructions that achieve guaranteed recovery, and apply the method to Vision Transformer representations to obtain early post‑activation separability certificates.
arXiv:2607. 07035v1 Announce Type: cross Abstract: The architecture of deep feedforward neural networks is ubiquitous in deep learning, either as a whole system or as a subnetwork of other architectures, and thus its mechanism is a key ingredient of the black box of neural networks.
arXiv:2606. 16028v1 Announce Type: new Abstract: Modern deep learning architectures are increasingly multi-task and multi-modal, using a pretrained foundation model combined with task-specific, fine-tuned models.
arXiv:2505.09716v3 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) generalisation is considered a hallmark of human and animal intelligence. To achieve OOD through composition, a sys...
arXiv:2609.39078v1 Announce Type: new Abstract: Representations are routinely used across machine learning, psychology, and neuroscience to draw inferences about the computations of biological and ar...
arXiv:2602. 01083v2 Announce Type: replace Abstract: Weight-space learning studies neural architectures that operate directly on the parameters of other neural networks.
arXiv:2310.16295v2 Announce Type: replace-cross Abstract: Neural network have achieved remarkable successes in many scientific fields. However, the interpretability of the neural network model is sti...
arXiv:2410. 06665v4 Announce Type: replace-cross Abstract: This paper explores the characterization of equivariant linear layers for representations of permutations and related groups.
CrossGMN introduces a graph metanetwork that processes a trained source network and an initialized target network simultaneously, enabling equivariant cross‑architecture weight‑space transformations. By preserving symmetry through cross‑network message passing, CrossGMN can refine target network initializations while remaining invariant to source permutations and equivariant to target permutations. Experiments demonstrate that CrossGMN accelerates knowledge distillation, transfers across datasets without retraining, and unifies compression from diverse source architectures into a common target architecture.