Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale.
arXiv:2603. 21014v2 Announce Type: replace Abstract: Mechanistic interpretability seeks to understand how Large Language Models (LLMs) represent and process information.
By Florent Draye, Vedant Palit, Abir Harrasse, Tung-Yu Wu, Jiarui Liu, Punya Syon Pandey, Roderick Wu, Chih-Hao Hsu, Terry Jingchen Zhang, Zhijing Jin, Bernhard Sch\"olkopf
We present a Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoder (SAE) for extracting cross-seed universal features from independently trained BERT models. Cross-seed feature universality is a fundamental challenge in mechanistic interpretability: because dictionary learning is non-convex, independently trained networks learn misaligned feature spaces, so apparently identical features may differ by random initialization.
MURANO is an open‑source framework that enables researchers to design, run, and reproduce mechanistic interpretability experiments on large language models. It unifies the five key stages—loading, recording, attribution, intervention, and evaluation—into composable pipeline steps that exchange named artifacts and use canonical addresses for interoperability. The authors demonstrate the framework with reproductions of existing studies and a sparse autoencoder case study, showing its practical applicability across disciplines.
By Alireza Bayat Makou, Emirhan B\"oge, Phu Gia Hoang, Federico Tiblias, Jingcheng Niu, Subhabrata Dutta, Richard Eckart de Castilho, Iryna Gurevych
arXiv:2606. 08496v1 Announce Type: cross Abstract: Although Sparse Autoencoders (SAEs) have mitigated the opacity of large language models (LLMs) by decomposing dense representations into sparse features, explaining these features still remains a central challenge.
By Jingyi He, Haiyan Zhao, Ruxue Shi, Yanguang Liu, Xin Wang, Fei Sun, Mengnan Du
arXiv:2606. 04928v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed across diverse applications, raising critical questions for governance, accountability, and data provenance.
By Fr\'ed\'eric Berdoz, Luca A. Lanzend\"orfer, Kaan Bayraktar, Roger Wattenhofer
arXiv:2606. 07524v1 Announce Type: cross Abstract: The explosive growth of large language models (LLMs) has created a heterogeneous and poorly documented ecosystem, making systematic model comparison increasingly important for provenance auditing, security analysis, and model selection.
By Zirui Wang, Yusen Hou, Shaofeng Liang, Bowen Tian, Yanlin Zhang, Wenshuo Chen, Yutao Yue
arXiv:2607. 22610v1 Announce Type: new Abstract: When a language model produces a response in a multi-turn conversation, which tokens from prior turns shaped that answer, and how did those dependencies propagate across prior turns?
By Jessica Tang, Shraddha Barke, Sharad Agarwal
arXiv:2609.07876v1 Announce Type: cross
Abstract: Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose residual stream contributions into linea...
By Arjun Patrawala, Jiahai Feng, Erik Jones, Jacob Steinhardt
SharedSAE demonstrates that a single sparse autoencoder can replace multiple model‑specific SAEs by using a shared dictionary with model‑specific encoder‑decoder pairs. It preserves activation magnitudes, normalizes only selection scores, and supports single‑model inference via model dropout. Trained on four 1B‑scale language models, SharedSAE retains 96.6% of the mean explained variance of dedicated SAEs, shows higher cross‑model latent correlations, and allows efficient adaptation of new models to the shared latent space.
By Daniil Ognev, C\'elian Vasson, Lijie Hu, Kentaro Inui, Benjamin Heinzerling
The paper presents an iterative framework that uses large language models (LLMs) to automatically extract interpretable, schema‑bound categorical features from unstructured text for use in tabular prediction models. A generator LLM proposes semantic definitions, an extractor LLM materializes the features, and a downstream tabular model evaluates their predictive performance, with error‑driven natural‑language feedback guiding the search. Across three public datasets, the error‑driven loop speeds up feature discovery up to three times and the resulting features outperform any subset when combined with TF‑IDF and dense embeddings, while also providing instance‑level interpretability through SHAP importance rankings and a semantic audit trail.
By Merwan Barlier, Blaz Skrlj
arXiv:2508. 17320v3 Announce Type: replace Abstract: Understanding the internal representations of large language models (LLMs) remains a central challenge for interpretability research.
By Yifei Yao, Hanrong Zhang, Mengnan Du