arXiv AI By Louie Hong Yao, Yuhao Li, Shengchao Liu

The Linear Representation Hypothesis Needs a Group Action

Read the original on arXiv AI →

The paper argues that the Linear Representation Hypothesis (LRH) should not be treated as a single claim but as a family of claims differentiated by how representations are considered equivalent. It highlights that different equivalence notions preserve different structures, leading to metrics, probes, and interventions that may actually test distinct hypotheses. By formalizing these ideas with group actions, the authors provide a framework that clarifies how assumptions vary across metrics, reading points, and analysis stages, and they apply it to audit common representation quantities and recent interpretability analyses.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 23

A Survey on the Linear Representation Hypothesis

arXiv:2609.22695v1 Announce Type: new Abstract: The term "linear representation hypothesis" (LRH) has appeared across diverse subfields of artificial intelligence, neuroscience, and cognitive science...

By Sewoong Lee, Marc E. Canby, Ikhyun Cho, Julia Hockenmaier
arXiv AI
Sep 25

Every Component Is a Lookup: One Linear Graph for Interaction, Composition and Attribution

The paper proposes that two architectural assumptions—(1) attention and MLPs share a key‑value form <phi(S)>U, and (2) components read from an additive residual stream—are sufficient to answer three interpretability questions: component interaction, information routing, and token attribution. By treating these selections as a computational graph, the authors develop Unpack, a backward attribution method that validates interaction scores, recovered routes, and token attribution against established tests across models ranging from 160M to 6.9B parameters. The study also shows that contribution and causal effect can differ, with a recognizable signature in how components change when a task is removed.

By Po-Kai Chen, Aske Plaat, Niki van Stein