arXiv Machine Learning
Aug 31

Towards a mathematical theory of superposition

The paper develops a mathematical theory of superposition in neural networks using frame theory and compressed sensing. It shows that a sparse binary vector of active features can be encoded by an overcomplete dictionary and recovered via a ReLU operation with a suitable bias. The authors prove recovery theorems for both random-support and worst-case support settings, providing high-probability guarantees for low-coherence dictionaries and a sharp criterion for sparsity levels, with explicit results for Gaussian random matrices and equiangular tight frames.

By Michael I. Ivanitskiy, John Jasper, Emily J. King, Dustin G. Mixon
arXiv Machine Learning
Jun 5

Expand Neurons, Not Parameters

arXiv:2510. 04500v3 Announce Type: replace Abstract: This work demonstrates how increasing the number of neurons in a network without increasing its total number of non-zero parameters improves performance.

By Linghao Kong, Inimai Subramanian, Yonadav Shavit, Micah Adler, Dan Alistarh, Nir Shavit