arXiv AI

The Linear Representation Hypothesis Needs a Group Action

The paper argues that the Linear Representation Hypothesis (LRH) should not be treated as a single claim but as a family of claims differentiated by how representations are considered equivalent. It highlights that different equivalence notions preserve different structures, leading to metrics, probes, and interventions that may actually test distinct hypotheses. By formalizing these ideas with group actions, the authors provide a framework that clarifies how assumptions vary across metrics, reading points, and analysis stages, and they apply it to audit common representation quantities and recent interpretability analyses.

arXiv AI
Sep 23

A Survey on the Linear Representation Hypothesis

arXiv:2609.22695v1 Announce Type: new Abstract: The term "linear representation hypothesis" (LRH) has appeared across diverse subfields of artificial intelligence, neuroscience, and cognitive science...

By Sewoong Lee, Marc E. Canby, Ikhyun Cho, Julia Hockenmaier
arXiv AI
Sep 25

Every Component Is a Lookup: One Linear Graph for Interaction, Composition and Attribution

The paper proposes that two architectural assumptions—(1) attention and MLPs share a key‑value form <phi(S)>U, and (2) components read from an additive residual stream—are sufficient to answer three interpretability questions: component interaction, information routing, and token attribution. By treating these selections as a computational graph, the authors develop Unpack, a backward attribution method that validates interaction scores, recovered routes, and token attribution against established tests across models ranging from 160M to 6.9B parameters. The study also shows that contribution and causal effect can differ, with a recognizable signature in how components change when a task is removed.

By Po-Kai Chen, Aske Plaat, Niki van Stein
arXiv AI
6d ago

The Linear Representation Hypothesis for Vision-Language-Action Models

The paper proposes a theoretical framework called the Linear Representation Hypothesis (LRH) for vision‑language‑action (VLA) models, extending the concept from large language models to systems where physical quantities of interest (QoI) evolve with the dynamics. It introduces a signature‑based formulation that unifies representations and policies, proving that future QoI evolution can be linearly probed from representations and that a generalized linear model for stochastic action chunks allows monotonic steering of QoI. The authors validate their theory with an explicit oracle representation in a planar control‑affine navigation experiment, demonstrating the predicted linear probing and steering mechanisms.

By Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han
arXiv AI
Jun 15

Actionable Interpretability Must Be Defined in Terms of Symmetries

arXiv:2601. 12913v4 Announce Type: replace Abstract: This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpretability can be formally tested or designed for.

By Pietro Barbiero, Mateo Espinosa Zarlenga, Francesco Giannini, Alberto Termine, Filippo Bonchi, Mateja Jamnik, Giuseppe Marra