arXiv Machine Learning

Directional Linear Separability of Neural Representations: Geometry and Transformations

The paper introduces the directional linear separability measure (D‑LSM) to quantify how much linear separability is preserved or improved by injective affine maps in neural networks. It characterizes the geometry supporting D‑LSM, proves its invariance under injective affine embeddings, and derives conditions for gated activations (ReLU, GELU, SiLU) to preserve and recover samples. Experiments validate the theoretical bounds, demonstrate affine‑tube constructions that achieve guaranteed recovery, and apply the method to Vision Transformer representations to obtain early post‑activation separability certificates.

arXiv Machine Learning
Jun 5

Separation Power of Equivariant Neural Networks

arXiv:2406. 08966v3 Announce Type: replace Abstract: The separation power of a machine learning model refers to its ability to distinguish between different inputs and is often used as a proxy for its expressivity.

By Marco Pacini, Xiaowen Dong, Bruno Lepri, Gabriele Santin
arXiv Machine Learning
Aug 31

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations

The paper introduces Mixture of Activations (MoA), a token‑adaptive feedforward network design that mixes multiple activation functions using lightweight gates while sharing linear projections. It also presents learnable activations (LA) as an input‑independent variant. The authors theoretically prove that MoA strictly surpasses both fixed‑activation FFNs and LA in expressive power, and empirically demonstrate that MoA achieves lower loss and better scaling on dense and MoE language models from 0.12 B to 2 B parameters with minimal overhead.

By Mingze Wang, Jinbo Wang, Yikuan Xia, Kai Shen, Shu Zhong
arXiv AI
Sep 15

L-Lipschitz Gershgorin ResNet Network

The paper introduces a method for constructing L-Lipschitz deep residual networks (ResNets) using a Linear Matrix Inequality (LMI) framework. By reformulating the ResNet architecture as a pseudo-tridiagonal LMI and applying the Gershgorin circle theorem, the authors derive closed‑form constraints on network parameters that guarantee Lipschitz continuity. The work also presents a compositional framework for handling recursive systems in hierarchical architectures, while noting that the Gershgorin-based approximations can over‑constrain the system, reducing expressive capacity.

By Marius F. R. Juston, William R. Norris, Dustin Nottage, Ahmet Soylemezoglu
arXiv AI
Jul 21

Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm

arXiv:2607. 16295v1 Announce Type: cross Abstract: Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning (SDL) methods, most prominently sparse autoencoders, as a central paradigm.

By Yiming Tang, Qinglin Qi, Zhaoqian Yao, Harshvardhan Saini, Dianbo Liu