arXiv:2607. 04800v1 Announce Type: new Abstract: Neural networks are thought to represent concepts as directions in their activation space, and superposition lets them encode more concepts than they have dimensions.
By Francisco Ferreira da Silva, Stefan Heimersheim
arXiv:2609.40127v1 Announce Type: cross
Abstract: Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keepin...
By Massimo Bini, Anders Christensen, Stephan Alaniz, Judah Goldfeder, Ole Winther, Yann LeCun, Ravid Shwartz-Ziv, Zeynep Akata
Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keeping the matrices dense, and thus efficient on standar...
The paper introduces Unmerge, an efficient machine unlearning algorithm that treats unlearning as the inverse of task arithmetic. By representing the forget component as a low‑rank basis at each layer, Unmerge optimizes three goals—matching the merged vector, suppressing leakage, and bounding correction size—to limit forget leakage and retain damage. Experiments on ResNet‑50, ViT‑S/16, and Llama‑3.2‑3B show significant performance gains over existing methods while maintaining privacy and feature‑distribution fidelity.
By Haoran Tang, Andrew Tan, Rajiv Khanna
arXiv:2604. 00208v2 Announce Type: replace Abstract: Comparing internal representations is a central goal in neuroscience and machine learning, but standard linear alignment metrics (Representational Similarity Analysis, Centered Kernel Alignment, and linear regression) are frequently applied to neural activity coordinates rather than on the underlying features.
By Sunny Liu, Habon Issa, Andr\'e Longon, Liv Gorton, Meenakshi Khosla, Alex Williams, David Klindt
arXiv:2606. 18538v1 Announce Type: new Abstract: One of the major difficulties in the mechanistic interpretability of neural networks is the occurrence of polysemanticity, which suggests that each neuron is typically responsible for multiple different tasks, impeding a clean interpretation of their function.
By Mriganka Basu Roy Chowdhury, Eric McLaughlin Weiner
Approximate machine unlearning seeks to remove the influence of a forget set from a trained model without full retraining. Existing gradient-based methods require data-dependent hyperparameter search,...
arXiv:2607. 11990v1 Announce Type: cross Abstract: Feedforward network (FFN) blocks account for a large fraction of the parameters and computation in Transformer architectures, yet their internal structure remains difficult to interpret due to the additive superposition induced by the residual stream.
By Johannes Knittel, Hanspeter Pfister
arXiv:2510. 04500v3 Announce Type: replace Abstract: This work demonstrates how increasing the number of neurons in a network without increasing its total number of non-zero parameters improves performance.
By Linghao Kong, Inimai Subramanian, Yonadav Shavit, Micah Adler, Dan Alistarh, Nir Shavit
arXiv:2606. 07414v1 Announce Type: new Abstract: Sparsity allows scaling model parameters without proportionally increasing computational cost.
By Simon Schug
arXiv:2607. 21366v1 Announce Type: cross Abstract: Deep neural networks encode complex representations, but deconstructing this internal knowledge remains a challenge.
By Hossein Mobahi, Peter L. Bartlett
arXiv:2609.06862v1 Announce Type: new
Abstract: Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons...
By Dai Shi, Xiaoyu Li, Andi Han, Jos\'e Miguel Hern\'andez-Lobato