Towards a mathematical theory of superposition
Read the original on arXiv Machine Learning →The paper develops a mathematical theory of superposition in neural networks using frame theory and compressed sensing. It shows that a sparse binary vector of active features can be encoded by an overcomplete dictionary and recovered via a ReLU operation with a suitable bias. The authors prove recovery theorems for both random-support and worst-case support settings, providing high-probability guarantees for low-coherence dictionaries and a sharp criterion for sparsity levels, with explicit results for Gaussian random matrices and equiangular tight frames.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.