arXiv:2609.22695v1 Announce Type: new
Abstract: The term "linear representation hypothesis" (LRH) has appeared across diverse subfields of artificial intelligence, neuroscience, and cognitive science...
By Sewoong Lee, Marc E. Canby, Ikhyun Cho, Julia Hockenmaier
The paper proposes that two architectural assumptions—(1) attention and MLPs share a key‑value form <phi(S)>U, and (2) components read from an additive residual stream—are sufficient to answer three interpretability questions: component interaction, information routing, and token attribution. By treating these selections as a computational graph, the authors develop Unpack, a backward attribution method that validates interaction scores, recovered routes, and token attribution against established tests across models ranging from 160M to 6.9B parameters. The study also shows that contribution and causal effect can differ, with a recognizable signature in how components change when a task is removed.
By Po-Kai Chen, Aske Plaat, Niki van Stein
arXiv:2609.37680v1 Announce Type: cross
Abstract: One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can...
By Sai Sumedh R. Hindupur, Hadas Orgad, Thomas Fel, Demba Ba
arXiv:2607. 17800v1 Announce Type: new Abstract: Representation is a central concept in modern machine learning, where it usually refers to internal encodings that support learning and generalization.
By Gilad Landau, Aviv Keren
arXiv:2607. 03598v1 Announce Type: cross Abstract: When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished project and it critiques the code; share a raw late-night line and it runs a wellness check.
By Alex Kwon
arXiv:2606. 30686v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) systems, built on pretrained vision-language models (VLMs), have shown rapidly improving performance on robot manipulation benchmarks.
By Taozhao Chen, Ian Manchester, Huaming Chen
arXiv:2606. 01092v1 Announce Type: cross Abstract: Supervised learning evaluates predictors through their input-output behavior.
By Vasileios Sevetlidis
arXiv:2608.29034v1 Announce Type: cross
Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, differ...
By Zhang Enyan, R. Thomas McCoy
arXiv:2607. 08843v1 Announce Type: new Abstract: In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space.
By William W. Yang, Andrew M. Saxe, Peter E. Latham
The paper proposes a theoretical framework called the Linear Representation Hypothesis (LRH) for vision‑language‑action (VLA) models, extending the concept from large language models to systems where physical quantities of interest (QoI) evolve with the dynamics. It introduces a signature‑based formulation that unifies representations and policies, proving that future QoI evolution can be linearly probed from representations and that a generalized linear model for stochastic action chunks allows monotonic steering of QoI. The authors validate their theory with an explicit oracle representation in a planar control‑affine navigation experiment, demonstrating the predicted linear probing and steering mechanisms.
By Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han
arXiv:2601. 12913v4 Announce Type: replace Abstract: This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpretability can be formally tested or designed for.
By Pietro Barbiero, Mateo Espinosa Zarlenga, Francesco Giannini, Alberto Termine, Filippo Bonchi, Mateja Jamnik, Giuseppe Marra
arXiv:2606. 07303v1 Announce Type: new Abstract: Representation learning is central to modern machine learning, enabling transitions from handcrafted features to learned embeddings, latent spaces, foundation models, world models, and digital twins.
By Jacques Raynal, Pierre Slangen, Elsa Raynal, Jacques Margerit