arXiv:2606. 07007v1 Announce Type: cross Abstract: We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs).
By Chenhao Zhang, Chris Lin, Su-In Lee
arXiv:2606. 06333v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) are widely used for mechanistic interpretability in large language models, yet their formulation assigns each latent feature a single decoder direction, implicitly assuming features to be one-dimensional.
By Seyed Arshan Dalili, Mehrdad Mahdavi
arXiv:2512. 07355v2 Announce Type: replace Abstract: Two traditions of interpretability have evolved side by side but seldom spoken to each other: Concept Bottleneck Models (CBMs), which prescribe what a concept should be, and Sparse Autoencoders (SAEs), which discover what concepts emerge.
By Alexandre Rocchi, Thomas Fel, Gianni Franchi
arXiv:2606. 13803v1 Announce Type: new Abstract: Enforcing functional inequality constraints such as monotonicity and convexity in neural networks is a fundamental challenge in many industrial and scientific applications.
By Ruben Wiedemann, Antoine Jacquier, Lukas Gonon
arXiv:2606. 02841v1 Announce Type: new Abstract: Deep neural networks learn representations where individual features often lack interpretable meaning; a single neuron may activate for scattered, unrelated inputs.
By Sigurd Gaukstad, Melvin Vaupel, Valdemar Karg{\aa}rd Olsen, Erik Hermansen, Benjamin Dunn
arXiv:2603. 04198v2 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) are widely used to extract human-interpretable features from neural network activations, but their learned features can vary substantially across random seeds and training choices.
By Piotr Jedryszek, Oliver M. Crook