arXiv AI

C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders

arXiv:2606. 30609v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse, human-understandable features, but scaling to large dictionaries exposes fundamental challenges.

arXiv Machine Learning
Sep 7

SharedSAE: One Feature Dictionary Across Language Models

SharedSAE demonstrates that a single sparse autoencoder can replace multiple model‑specific SAEs by using a shared dictionary with model‑specific encoder‑decoder pairs. It preserves activation magnitudes, normalizes only selection scores, and supports single‑model inference via model dropout. Trained on four 1B‑scale language models, SharedSAE retains 96.6% of the mean explained variance of dedicated SAEs, shows higher cross‑model latent correlations, and allows efficient adaptation of new models to the shared latent space.

By Daniil Ognev, C\'elian Vasson, Lijie Hu, Kentaro Inui, Benjamin Heinzerling
arXiv Machine Learning
Jun 18

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability

arXiv:2606. 18383v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable features from language models (LMs), yet a central question remains: when can an SAE-based explanation be treated as a faithful view of an underlying frozen LM We study this through a post-hoc generalization framework that certifies the LM via a sparse proxy, obtained by replacing a native hidden activation with its pretrained SAE reconstruction.

By Dibyanayan Bandyopadhyay, Asif Ekbal
arXiv Machine Learning
Sep 3

Persistent Sparse Autoencoders: Learning Feature-Specific Timescales in Language Model Representations

Persistent Sparse Autoencoders (Persistent SAEs) extend standard sparse autoencoders by learning a persistence coefficient for each feature, enabling the model to capture feature‑specific timescales from reconstruction alone. The study shows that these persistent features maintain competitive reconstruction quality while distinguishing between short‑timescale, locally interpretable features and long‑timescale, context‑accumulating features. In a prompt‑injection monitoring case study, slow features were found to preserve injection‑related signals and remain causally effective over long contexts.

By Haoyan Luo, Mateo Espinosa Zarlenga, Mateja Jamnik