arXiv:2603. 04198v2 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) are widely used to extract human-interpretable features from neural network activations, but their learned features can vary substantially across random seeds and training choices.
By Piotr Jedryszek, Oliver M. Crook
arXiv:2606. 27941v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) provide useful decompositions of Transformer residual streams, but their learned features are usually named post hoc rather than directly connected to the Transformer's token vocabulary.
By Kairui Zhang, Ziwen Yu, Zahraa S. Abdallah, Martha Lewis
arXiv:2605. 18629v2 Announce Type: replace Abstract: Sparse autoencoders (SAEs) are one of the main methods to interpret the inner workings of deep neural networks (DNNs), decomposing activations into higher-dimensional features.
By Micha{\l} Brzozowski, Neo Christopher Chung
arXiv:2606. 27321v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation models, decomposing their polysemantic activations into a larger set of sparse, more monosemantic features.
By Nathana\"el Jacquier, Maria Vakalopoulou, Mahdi S. Hosseini
arXiv:2607. 17770v1 Announce Type: cross Abstract: Within Explainable Artificial Intelligence, mechanistic interpretability uses Sparse Autoencoders (SAEs) to extract more interpretable features from neural representations.
By Katarzyna Filus, Sebastian Pokuci\'nski
The standard way to compare two text embeddings is cosine similarity. Scattered studies report that a different metric does better, but never pin down the geometric condition that decides when, or why.
arXiv:2504. 16318v3 Announce Type: replace Abstract: Cosine similarity is a standard comparison rule for learned representations in information retrieval, natural language processing, computer vision, and multimodal learning.
By Kisung You
arXiv:2607. 08605v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept.
By Weiduo Liao, Yunqiao Yang, Ying Wei
arXiv:2606. 30609v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse, human-understandable features, but scaling to large dictionaries exposes fundamental challenges.
By Haoran Jin, Xiting Wang, Shijie Ren, Hong Xie, Defu Lian
arXiv:2602. 19393v2 Announce Type: replace Abstract: Steck, Ekanadham, and Kallus [arXiv:2403.
By Taha Bouhsine
arXiv:2607. 17568v1 Announce Type: cross Abstract: Structured pruning compresses large language models (LLMs) by removing whole computational units, such as attention heads and feed-forward (FFN) channel groups.
By Zhiren Gong, Zihao Zeng, Zijie Wang, Tiantong Wang, Chau Yuen, Wei Yang Bryan Lim
arXiv:2608. 14922v1 Announce Type: cross Abstract: Mechanistic interpretability has recently expanded to Vision Transformers (ViTs), with Sparse Autoencoders (SAEs) increasingly used as post-hoc tools to decompose internal representations into sparse and more interpretable features.
By Philip H. Lee, Parth Padalkar