arXiv:2606. 18538v1 Announce Type: new Abstract: One of the major difficulties in the mechanistic interpretability of neural networks is the occurrence of polysemanticity, which suggests that each neuron is typically responsible for multiple different tasks, impeding a clean interpretation of their function.
By Mriganka Basu Roy Chowdhury, Eric McLaughlin Weiner
arXiv:2609.06862v1 Announce Type: new
Abstract: Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons...
By Dai Shi, Xiaoyu Li, Andi Han, Jos\'e Miguel Hern\'andez-Lobato
The paper develops a mathematical theory of superposition in neural networks using frame theory and compressed sensing. It shows that a sparse binary vector of active features can be encoded by an overcomplete dictionary and recovered via a ReLU operation with a suitable bias. The authors prove recovery theorems for both random-support and worst-case support settings, providing high-probability guarantees for low-coherence dictionaries and a sharp criterion for sparsity levels, with explicit results for Gaussian random matrices and equiangular tight frames.
By Michael I. Ivanitskiy, John Jasper, Emily J. King, Dustin G. Mixon
arXiv:2503. 07325v2 Announce Type: replace Abstract: Understanding and certifying the behavior of modern deep neural networks remains a fundamental challenge in reliable machine learning.
By Khoat Than, Dat Phan
arXiv:2510. 04500v3 Announce Type: replace Abstract: This work demonstrates how increasing the number of neurons in a network without increasing its total number of non-zero parameters improves performance.
By Linghao Kong, Inimai Subramanian, Yonadav Shavit, Micah Adler, Dan Alistarh, Nir Shavit
arXiv:2511. 09432v2 Announce Type: replace Abstract: Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity.
By Ege Erdogan, Ana Lucic
arXiv:2608.20812v1 Announce Type: new
Abstract: We develop constructive approximation and learning guarantees for shallow neural models with infinite-dimensional inputs observed through finitely many...
By Pablo M. Bern\'a, Antonio Falc\'o, Diego Mond\'ejar
arXiv:2606. 02385v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) have found success parsing neural representations into interpretable concepts, providing a basis for understanding and control.
By William Dorrell
arXiv:2609. 19077v1 Announce Type: cross Abstract: Formal explainability provides mathematically grounded justifications for individual predictions.
By Frederic Koriche, Jean-Marie Lagniez, Chi Tran
arXiv:2606. 16028v1 Announce Type: new Abstract: Modern deep learning architectures are increasingly multi-task and multi-modal, using a pretrained foundation model combined with task-specific, fine-tuned models.
By Thomas Dittrich, Oliver Potocki, Philipp Grohs
arXiv:2606. 07007v1 Announce Type: cross Abstract: We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs).
By Chenhao Zhang, Chris Lin, Su-In Lee
arXiv:2604. 00208v2 Announce Type: replace Abstract: Comparing internal representations is a central goal in neuroscience and machine learning, but standard linear alignment metrics (Representational Similarity Analysis, Centered Kernel Alignment, and linear regression) are frequently applied to neural activity coordinates rather than on the underlying features.
By Sunny Liu, Habon Issa, Andr\'e Longon, Liv Gorton, Meenakshi Khosla, Alex Williams, David Klindt