SharedSAE demonstrates that a single sparse autoencoder can replace multiple model‑specific SAEs by using a shared dictionary with model‑specific encoder‑decoder pairs. It preserves activation magnitudes, normalizes only selection scores, and supports single‑model inference via model dropout. Trained on four 1B‑scale language models, SharedSAE retains 96.6% of the mean explained variance of dedicated SAEs, shows higher cross‑model latent correlations, and allows efficient adaptation of new models to the shared latent space.
By Daniil Ognev, C\'elian Vasson, Lijie Hu, Kentaro Inui, Benjamin Heinzerling
arXiv:2602.01695v2 Announce Type: replace
Abstract: Latent reasoning reduces the token-generation cost of chain-of-thought reasoning by replacing explicit intermediate tokens with continuous latent t...
By Yadong Wang, Haodong Chen, Yu Tian, Chuanxing Geng, Dong Liang, Xiang Chen
arXiv:2606. 30609v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse, human-understandable features, but scaling to large dictionaries exposes fundamental challenges.
By Haoran Jin, Xiting Wang, Shijie Ren, Hong Xie, Defu Lian
arXiv:2606. 10029v1 Announce Type: cross Abstract: Language models increasingly serve as the backbone of text-to-speech (TTS) systems, yet we understand little about the representations they build when text and generated speech tokens share a single residual stream.
By Nikita Koriagin, Georgii Aparin, Nikita Balagansky, Daniil Gavrilov
Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale.
arXiv:2608.28806v1 Announce Type: new
Abstract: Sparse autoencoders (SAEs) disentangle model activations into interpretable features and are widely used for steering large language models. Most exist...
By Yutian Liu, Xu Wang, Difan Zou
arXiv:2508. 17320v3 Announce Type: replace Abstract: Understanding the internal representations of large language models (LLMs) remains a central challenge for interpretability research.
By Yifei Yao, Hanrong Zhang, Mengnan Du
arXiv:2607. 00023v1 Announce Type: cross Abstract: Dense sentence embeddings are fundamental to modern Retrieval-Augmented Generation (RAG) systems but suffer from a lack of interpretability due to feature superposition.
By Wonseok Shin, Songkuk Kim
The paper introduces the Superposed Latent Autoencoder (SLAE), a method that stores multiple wide latent representations together by superposing them into a single memory tensor using learned codes and randomized keys. SLAE eliminates the need for tight dimensional bottlenecks, achieving up to 56% lower reconstruction error on datasets such as CIFAR-10/100 and SVHN while maintaining the same storage budget. The approach also boosts downstream classification performance by up to 16.79 percentage points, demonstrating that wide representations can be effectively compressed through structured interference rather than dimensional reduction.
By Quanling Zhao, Jiaying Yang, Tianqi Zhang, Ziyang Hao, Fatemeh Asgarinejad, Flavio Ponzina, Tajana Rosing
arXiv:2606. 14990v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are standard tools for mechanistic interpretability, but current SAE families are constrained by fixed encoder nonlinearities such as ReLU, JumpReLU, and TopK.
By Naiyu Yin, Yue Yu
arXiv:2605.12225v3 Announce Type: replace
Abstract: While deep transformer-based models have advanced rapidly, their internal mechanisms remain largely a mystery. Recent work has prioritized understa...
By Dan Pluth, Zachary Nicholas Houghton, Yu Zhou, Vijay K. Gurbani
The paper introduces DSPA, a dynamic sparse autoencoder (SAE) steering technique that aligns language model outputs with user preferences during inference, avoiding costly weight updates. DSPA constructs a conditional-difference map from preference triples to adjust token-active latents, improving MT‑Bench scores and matching AlpacaEval performance on models like Gemma‑2 and Qwen3 while preserving accuracy. It demonstrates robustness with limited preference data, outperforms the two‑stage RAHF‑SCIT pipeline in FLOPs, and reveals that preference directions are largely driven by discourse and stylistic cues.
By James Wedgwood, Aashiq Muhamed, Mona T. Diab, Virginia Smith