A Survey on the Linear Representation Hypothesis
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper argues that the Linear Representation Hypothesis (LRH) should not be treated as a single claim but as a family of claims differentiated by how representations are considered equivalent. It highlights that different equivalence notions preserve different structures, leading to metrics, probes, and interventions that may actually test distinct hypotheses. By formalizing these ideas with group actions, the authors provide a framework that clarifies how assumptions vary across metrics, reading points, and analysis stages, and they apply it to audit common representation quantities and recent interpretability analyses.
arXiv:2607. 17800v1 Announce Type: new Abstract: Representation is a central concept in modern machine learning, where it usually refers to internal encodings that support learning and generalization.
arXiv:2508. 11214v2 Announce Type: replace-cross Abstract: Explanations of cognitive behavior often appeal to computations over representations.
arXiv:2609.06862v1 Announce Type: new Abstract: Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons...
arXiv:2607. 08843v1 Announce Type: new Abstract: In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space.
arXiv:2608.29034v1 Announce Type: cross Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, differ...