arXiv:2603. 04198v2 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) are widely used to extract human-interpretable features from neural network activations, but their learned features can vary substantially across random seeds and training choices.
By Piotr Jedryszek, Oliver M. Crook
arXiv:2606. 03002v1 Announce Type: cross Abstract: Quantization is a standard path to deploying large language models, and a quantized model is typically judged acceptable when its perplexity or downstream accuracy stays close to the full-precision original.
By Evan Duan
arXiv:2506.12576v3 Announce Type: replace
Abstract: Sparse autoencoders (SAEs) can enable inference-time topic steering by modifying latent feature activations, but existing steering methods often fa...
By Ananya Joshi, Celia Cintas, Skyler Speakman
arXiv:2606. 14990v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are standard tools for mechanistic interpretability, but current SAE families are constrained by fixed encoder nonlinearities such as ReLU, JumpReLU, and TopK.
By Naiyu Yin, Yue Yu
arXiv:2608.28806v1 Announce Type: new
Abstract: Sparse autoencoders (SAEs) disentangle model activations into interpretable features and are widely used for steering large language models. Most exist...
By Yutian Liu, Xu Wang, Difan Zou
arXiv:2606. 08365v1 Announce Type: cross Abstract: Sparse autoencoder (SAE) features are increasingly used to steer language models, but feature steering is rarely clean: the same intervention can behave inconsistently across contexts and perturb unrelated features.
By Evan Duan
arXiv:2607. 19386v1 Announce Type: new Abstract: Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores.
By Sinie van der Ben, Neele Roch, Anna Hedstr\"om, Mennatallah El-Assady
arXiv:2606. 30609v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse, human-understandable features, but scaling to large dictionaries exposes fundamental challenges.
By Haoran Jin, Xiting Wang, Shijie Ren, Hong Xie, Defu Lian
arXiv:2606. 18383v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable features from language models (LMs), yet a central question remains: when can an SAE-based explanation be treated as a faithful view of an underlying frozen LM We study this through a post-hoc generalization framework that certifies the LM via a sparse proxy, obtained by replacing a native hidden activation with its pretrained SAE reconstruction.
By Dibyanayan Bandyopadhyay, Asif Ekbal
arXiv:2605.16339v2 Announce Type: replace
Abstract: Preference learning in large language models relies on reward models as proxies for human judgment. However, these models frequently exhibit prefer...
By Shunchang Liu, Xin Chen, Belen Martin Urcelay, Francesco Croce
arXiv:2606. 12138v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on whether the learned features are reproducible across training runs.
By Gleb Gerasimov, Timofei Rusalev, Nikita Balagansky, Daniil Laptev, Vadim Kurochkin, Daniil Gavrilov
arXiv:2505. 16077v2 Announce Type: replace Abstract: Sparse autoencoders (SAEs) are used to decompose neural network activations into human-interpretable features.
By Soham Gadgil, Chris Lin, Su-In Lee