arXiv:2607. 04800v1 Announce Type: new Abstract: Neural networks are thought to represent concepts as directions in their activation space, and superposition lets them encode more concepts than they have dimensions.
By Francisco Ferreira da Silva, Stefan Heimersheim
arXiv:2609.06862v1 Announce Type: new
Abstract: Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons...
By Dai Shi, Xiaoyu Li, Andi Han, Jos\'e Miguel Hern\'andez-Lobato
arXiv:2609.09556v1 Announce Type: new
Abstract: Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear access...
By Enrico Vompa
arXiv:2510. 04500v3 Announce Type: replace Abstract: This work demonstrates how increasing the number of neurons in a network without increasing its total number of non-zero parameters improves performance.
By Linghao Kong, Inimai Subramanian, Yonadav Shavit, Micah Adler, Dan Alistarh, Nir Shavit
arXiv:2511. 09432v2 Announce Type: replace Abstract: Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity.
By Ege Erdogan, Ana Lucic
arXiv:2604. 00208v2 Announce Type: replace Abstract: Comparing internal representations is a central goal in neuroscience and machine learning, but standard linear alignment metrics (Representational Similarity Analysis, Centered Kernel Alignment, and linear regression) are frequently applied to neural activity coordinates rather than on the underlying features.
By Sunny Liu, Habon Issa, Andr\'e Longon, Liv Gorton, Meenakshi Khosla, Alex Williams, David Klindt
arXiv:2607. 08605v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept.
By Weiduo Liao, Yunqiao Yang, Ying Wei
arXiv:2508. 16560v4 Announce Type: replace-cross Abstract: Sparse Autoencoders (SAEs) extract features from LLM internal activations, meant to correspond to interpretable concepts.
By David Chanin, Adri\`a Garriga-Alonso
arXiv:2606. 07007v1 Announce Type: cross Abstract: We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs).
By Chenhao Zhang, Chris Lin, Su-In Lee
arXiv:2606. 12138v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on whether the learned features are reproducible across training runs.
By Gleb Gerasimov, Timofei Rusalev, Nikita Balagansky, Daniil Laptev, Vadim Kurochkin, Daniil Gavrilov
arXiv:2606. 09940v1 Announce Type: cross Abstract: Dictionary learning methods like Sparse Autoencoders (SAEs) and crosscoders attempt to explain a model by decomposing its activations into independent features.
By Dmitry Manning-Coe, Thomas Read, Anna Soligo, Oliver Clive-Griffin, Chun-Hei Yip, Rajashree Agrawal, Jason Gross
arXiv:2602. 04078v2 Announce Type: replace-cross Abstract: Deep learning has achieved remarkable success across a wide range of domains, significantly expanding the frontiers of what is achievable in artificial intelligence.
By R\'ois\'in Luo