Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders
arXiv:2608. 11197v1 Announce Type: new Abstract: Shani et al.
The study investigates how part‑of‑speech (PoS) categories are represented in the latent space of Sparse AutoEncoders (SAEs) applied to language models. Results show that PoS distinctions can be reliably recovered from SAE activations, but the mapping is not one‑to‑one; instead, PoS categories are supported by compact, distributed groups of sparse latents that vary across tags and remain stable on held‑out data. The findings suggest that SAEs encode morpho‑syntactic information in a distributed, category‑dependent manner rather than through isolated grammatical features.
arXiv:2608. 11197v1 Announce Type: new Abstract: Shani et al.
The paper compares latent representations in Selective State Space Models (SSMs) like Mamba and Transformers such as Pythia using Sparse Autoencoders. Across a 10‑million token corpus, 99.98% of Mamba features align closely with Pythia’s, supporting the Universality Hypothesis that core semantic representations are similar across architectures. A tiny 0.02% of features diverge, with Mamba’s recurrent bottleneck causing it to compress syntactic anomalies into polysemantic neurons, whereas Pythia’s attention can isolate distinct formatting edge‑cases.
The paper introduces LLM-Microscope, a toolkit for measuring how Large Language Models encode contextual information at the token level. It shows that seemingly minor tokens—such as determiners, stopwords, and punctuation—carry surprisingly high contextual weight, and removing them degrades performance on benchmarks like MMLU and BABILong-4k. The study also finds a strong link between contextualization and linearity, indicating that the transformation between layers can be approximated by a single linear mapping when tokens are well contextualized.
arXiv:2609.07876v1 Announce Type: cross Abstract: Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose residual stream contributions into linea...
The paper studies how transformer representations evolve across layers by examining the intrinsic dimensionality (ID) of token embeddings and their neighborhood structures. It finds that closed‑class tokens expand and collapse earlier than open‑class tokens, and that these changes are linked to shifts in local geometry. The authors compare encoder and decoder models, showing distinct layer‑wise behaviors, and demonstrate that geometric features alone can predict a token’s part‑of‑speech and reveal how semantic content changes across layers.
Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale.
arXiv:2609.00416v1 Announce Type: new Abstract: Probing studies have established that syntactic information is decodable in early and middle transformer layers, but what happens to that information i...
SharedSAE demonstrates that a single sparse autoencoder can replace multiple model‑specific SAEs by using a shared dictionary with model‑specific encoder‑decoder pairs. It preserves activation magnitudes, normalizes only selection scores, and supports single‑model inference via model dropout. Trained on four 1B‑scale language models, SharedSAE retains 96.6% of the mean explained variance of dedicated SAEs, shows higher cross‑model latent correlations, and allows efficient adaptation of new models to the shared latent space.
arXiv:2607. 17117v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) decompose language model activations into sparse features, but standard SAEs encode each token independently and do not expose information that persists across a sequence.
arXiv:2604. 02029v2 Announce Type: replace Abstract: Latent space is rapidly emerging as a native substrate for language-based models.
The study probes Gemma‑2‑9B‑IT with Sparse Autoencoders across English, Hebrew, and Russian to examine how multilingual LLMs handle informal register. By using a dataset of polysemous terms that appear in literal and informal contexts, the authors isolate pragmatic register processing from lexical cues. They discover a small, robust cross‑linguistic core that forms an informal register subspace, which becomes clearer in deeper layers and can causally shift output formality across all tested languages, even transferring zero‑shot to six unseen languages.
arXiv:2606.22473v2 Announce Type: replace-cross Abstract: Speech language models (SLMs) increasingly combine speech and text, often by interleaving their tokens within a single sequence. Yet how thes...