The paper introduces the directional linear separability measure (D‑LSM) to quantify how much linear separability is preserved or improved by injective affine maps in neural networks. It characterizes the geometry supporting D‑LSM, proves its invariance under injective affine embeddings, and derives conditions for gated activations (ReLU, GELU, SiLU) to preserve and recover samples. Experiments validate the theoretical bounds, demonstrate affine‑tube constructions that achieve guaranteed recovery, and apply the method to Vision Transformer representations to obtain early post‑activation separability certificates.
By Yi Wei, Xuan Qi, Suorong Yang, Furao Shen
arXiv:2311. 02960v5 Announce Type: replace Abstract: Over the past decade, deep learning has proven to be a highly effective tool for learning meaningful features from raw data.
By Peng Wang, Xiao Li, Can Yaras, Zhihui Zhu, Laura Balzano, Wei Hu, Qing Qu
arXiv:2608.30028v1 Announce Type: new
Abstract: This paper introduces a family of multiclass linear Perceptron classifiers with a multiplicative margin mechanism (MMPerc), as an alternative to standa...
By Dmitri Rachkovskij, Evgeny Osipov, Olexander Volkov, Daswin De Silva, Denis Kleyko
arXiv:2602. 24264v2 Announce Type: replace-cross Abstract: Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems.
By Arnas Uselis, Andrea Dittadi, Seong Joon Oh
arXiv:2606. 16028v1 Announce Type: new Abstract: Modern deep learning architectures are increasingly multi-task and multi-modal, using a pretrained foundation model combined with task-specific, fine-tuned models.
By Thomas Dittrich, Oliver Potocki, Philipp Grohs
arXiv:2607. 18930v1 Announce Type: cross Abstract: The Universal Approximation Theorem states that a neural network with a single hidden layer is sufficient to approximate any continuous univariate function on a compact domain to arbitrary error.
By Anuragine S A, Prem Jagadeesan