arXiv Machine Learning

Separation Power of Equivariant Neural Networks

arXiv:2406. 08966v3 Announce Type: replace Abstract: The separation power of a machine learning model refers to its ability to distinguish between different inputs and is often used as a proxy for its expressivity.

Hugging Face Trending Papers
Jun 1

Expressivity of congruence-based architectures for DNNs on positive-definite matrices

This work studies neural architectures for classifying symmetric positive-definite matrices, focusing on congruence-like layers, in which the input matrix is multiplied on the left and right by a (possibly rectangular) weight matrix $W$ and its transpose. Such layers lie at the core of the celebrated SPDNet and have also been employed independently for dimensionality reduction on positive-definite data.

arXiv Machine Learning
Sep 22

Directional Linear Separability of Neural Representations: Geometry and Transformations

The paper introduces the directional linear separability measure (D‑LSM) to quantify how much linear separability is preserved or improved by injective affine maps in neural networks. It characterizes the geometry supporting D‑LSM, proves its invariance under injective affine embeddings, and derives conditions for gated activations (ReLU, GELU, SiLU) to preserve and recover samples. Experiments validate the theoretical bounds, demonstrate affine‑tube constructions that achieve guaranteed recovery, and apply the method to Vision Transformer representations to obtain early post‑activation separability certificates.

By Yi Wei, Xuan Qi, Suorong Yang, Furao Shen
arXiv AI
Jul 9

On the Principles of Deep Feedforward ReLU Networks

arXiv:2607. 07035v1 Announce Type: cross Abstract: The architecture of deep feedforward neural networks is ubiquitous in deep learning, either as a whole system or as a subnetwork of other architectures, and thus its mechanism is a key ingredient of the black box of neural networks.

By Changcun Huang
arXiv Machine Learning
1d ago

CrossGMN: Graph Metanetworks for Cross-Architecture Weight-Space Transformations

CrossGMN introduces a graph metanetwork that processes a trained source network and an initialized target network simultaneously, enabling equivariant cross‑architecture weight‑space transformations. By preserving symmetry through cross‑network message passing, CrossGMN can refine target network initializations while remaining invariant to source permutations and equivariant to target permutations. Experiments demonstrate that CrossGMN accelerates knowledge distillation, transfers across datasets without retraining, and unifies compression from diverse source architectures into a common target architecture.

By Adir Dayan, Yam Eitan, Haggai Maron