arXiv Machine Learning

A Geometric Measure of Linear Separability for Neural Representations

arXiv:2606. 08721v1 Announce Type: new Abstract: Modern neural classifiers commonly rely on linear readouts, yet predictive metrics alone do not characterize the class-wise geometry of the representations on which such readouts operate.

arXiv Machine Learning
Sep 22

Directional Linear Separability of Neural Representations: Geometry and Transformations

The paper introduces the directional linear separability measure (D‑LSM) to quantify how much linear separability is preserved or improved by injective affine maps in neural networks. It characterizes the geometry supporting D‑LSM, proves its invariance under injective affine embeddings, and derives conditions for gated activations (ReLU, GELU, SiLU) to preserve and recover samples. Experiments validate the theoretical bounds, demonstrate affine‑tube constructions that achieve guaranteed recovery, and apply the method to Vision Transformer representations to obtain early post‑activation separability certificates.

By Yi Wei, Xuan Qi, Suorong Yang, Furao Shen
arXiv Machine Learning
Sep 25

Pointwise Generalization in Deep Neural Networks

The paper introduces a pointwise generalization theory for fully connected deep neural networks, using a pointwise Riemannian Dimension derived from eigenvalues of learned feature representations across layers. This framework provides hypothesis-dependent, representation-aware generalization bounds that are significantly tighter than traditional size- or norm-based approaches, both theoretically and experimentally. The authors analytically identify structural properties that explain deep networks’ tractability and empirically show that the pointwise Riemannian Dimension captures feature compression, over‑parameterization effects, and optimizer bias.

By Shaojie Li, Yunbei Xu
arXiv Machine Learning
Sep 17

NObSP: Functional Decomposition of Neural Networks via Oblique Subspace Projections

NObSP (Nonlinear Oblique Subspace Projections) is a framework that decomposes neural network predictions into explicit per‑feature contribution functions and an interaction residual, leveraging the linear final layer and oblique projections to avoid double counting when feature subspaces overlap. It connects to functional ANOVA and the Kolmogorov‑Arnold representation theorem, and introduces an efficient partial regression algorithm for out‑of‑sample evaluation. For convolutional networks, NObSP‑CAM generates class activation maps without backward passes after a single calibration, and experiments on tabular and vision datasets show faithfulness comparable to established attribution methods, with high function reproduction scores and improved class purity on TinyImageNet.

By Alexander Caicedo, V\'ictor De La Hoz, Santiago Alf\'erez