High-probability guarantees for linear accessibility in feature superposition
Read the original on arXiv Statistics ML →The Flow has not summarised this story yet — read it at arXiv Statistics ML.
The Flow has not summarised this story yet — read it at arXiv Statistics ML.
arXiv:2606. 18538v1 Announce Type: new Abstract: One of the major difficulties in the mechanistic interpretability of neural networks is the occurrence of polysemanticity, which suggests that each neuron is typically responsible for multiple different tasks, impeding a clean interpretation of their function.
arXiv:2609.06862v1 Announce Type: new Abstract: Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons...
The paper develops a mathematical theory of superposition in neural networks using frame theory and compressed sensing. It shows that a sparse binary vector of active features can be encoded by an overcomplete dictionary and recovered via a ReLU operation with a suitable bias. The authors prove recovery theorems for both random-support and worst-case support settings, providing high-probability guarantees for low-coherence dictionaries and a sharp criterion for sparsity levels, with explicit results for Gaussian random matrices and equiangular tight frames.
arXiv:2503. 07325v2 Announce Type: replace Abstract: Understanding and certifying the behavior of modern deep neural networks remains a fundamental challenge in reliable machine learning.
arXiv:2510. 04500v3 Announce Type: replace Abstract: This work demonstrates how increasing the number of neurons in a network without increasing its total number of non-zero parameters improves performance.
arXiv:2511. 09432v2 Announce Type: replace Abstract: Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity.