Language Model Circuits Are Sparse in the Neuron Basis
arXiv:2601. 22594v2 Announce Type: replace-cross Abstract: The high-level concepts that a neural network uses to perform computation need not be aligned to individual neurons (Smolensky, 1986).
arXiv:2603. 21014v2 Announce Type: replace Abstract: Mechanistic interpretability seeks to understand how Large Language Models (LLMs) represent and process information.
arXiv:2601. 22594v2 Announce Type: replace-cross Abstract: The high-level concepts that a neural network uses to perform computation need not be aligned to individual neurons (Smolensky, 1986).
arXiv:2606. 07524v1 Announce Type: cross Abstract: The explosive growth of large language models (LLMs) has created a heterogeneous and poorly documented ecosystem, making systematic model comparison increasingly important for provenance auditing, security analysis, and model selection.
arXiv:2606. 28548v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have become a useful tool for extracting interpretable features in language models.
arXiv:2606. 15796v1 Announce Type: cross Abstract: Mechanistic interpretability seeks to explain neural network behavior by decomposing model computations into interpretable features and circuits.
arXiv:2606. 11660v1 Announce Type: new Abstract: Data attribution is a promising field in interpretability that aims to explain model behavior through the influence of its training data, with applications including debugging undesirable model behavior and training dataset curation.
arXiv:2607. 20652v1 Announce Type: cross Abstract: Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams.
MURANO is an open‑source framework that enables researchers to design, run, and reproduce mechanistic interpretability experiments on large language models. It unifies the five key stages—loading, recording, attribution, intervention, and evaluation—into composable pipeline steps that exchange named artifacts and use canonical addresses for interoperability. The authors demonstrate the framework with reproductions of existing studies and a sparse autoencoder case study, showing its practical applicability across disciplines.
arXiv:2608. 02879v1 Announce Type: new Abstract: The widespread adoption of proprietary Large Language Models (LLMs) accessed strictly through closed APIs has created a critical challenge for responsible deployment: a fundamental lack of interpretability.
arXiv:2607. 20477v1 Announce Type: new Abstract: {\em Text-Attributed Graphs} (TAGs) have emerged as an expressive data model for integrating graph topology with rich textual semantics.
arXiv:2608. 07594v1 Announce Type: cross Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish.
arXiv:2606. 11722v1 Announce Type: cross Abstract: Finding interpretable directions in language-model representations is critical for understanding and controlling model behavior.
arXiv:2606. 31166v1 Announce Type: cross Abstract: Text-attributed graphs (TAGs), where each node carries a natural language description, require models to jointly reason over text and graph topology.