The paper argues that the Linear Representation Hypothesis (LRH) should not be treated as a single claim but as a family of claims differentiated by how representations are considered equivalent. It highlights that different equivalence notions preserve different structures, leading to metrics, probes, and interventions that may actually test distinct hypotheses. By formalizing these ideas with group actions, the authors provide a framework that clarifies how assumptions vary across metrics, reading points, and analysis stages, and they apply it to audit common representation quantities and recent interpretability analyses.
By Louie Hong Yao, Yuhao Li, Shengchao Liu
arXiv:2607. 17800v1 Announce Type: new Abstract: Representation is a central concept in modern machine learning, where it usually refers to internal encodings that support learning and generalization.
By Gilad Landau, Aviv Keren
arXiv:2508. 11214v2 Announce Type: replace-cross Abstract: Explanations of cognitive behavior often appeal to computations over representations.
By Atticus Geiger, Jacqueline Harding, Thomas Icard
arXiv:2609.06862v1 Announce Type: new
Abstract: Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons...
By Dai Shi, Xiaoyu Li, Andi Han, Jos\'e Miguel Hern\'andez-Lobato
arXiv:2607. 08843v1 Announce Type: new Abstract: In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space.
By William W. Yang, Andrew M. Saxe, Peter E. Latham
arXiv:2608.29034v1 Announce Type: cross
Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, differ...
By Zhang Enyan, R. Thomas McCoy
arXiv:2602. 24264v2 Announce Type: replace-cross Abstract: Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems.
By Arnas Uselis, Andrea Dittadi, Seong Joon Oh
arXiv:2608. 08159v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly reported to exhibit human-like neural and cognitive signatures, including concept cells, mental number lines, and cognitive maps.
By Yuqi Wu, Shengming Zhao, Jie Chen
arXiv:2606. 07303v1 Announce Type: new Abstract: Representation learning is central to modern machine learning, enabling transitions from handcrafted features to learned embeddings, latent spaces, foundation models, world models, and digital twins.
By Jacques Raynal, Pierre Slangen, Elsa Raynal, Jacques Margerit
arXiv:2606. 14512v1 Announce Type: cross Abstract: The recent successes of neural networks producing human-like language have caused significant stir in cognitive science, with many researchers arguing that classical puzzles about human cognition and challenges to artificial intelligence are being solved by neural networks.
By Michael Goodale, Salvador Mascarenhas
arXiv:2604.27927v2 Announce Type: replace
Abstract: We introduce a framework called LAPITHS (Language model Analysis through Paradigm grounded Interpretations of Theses about Human likenesS) and use...
By Matteo Da Pelo, Alessio Donvito, Claudio Frongia, Pietro Salis, Antonio Lieto
arXiv:2609.38998v1 Announce Type: new
Abstract: A foundational principle of connectionism is that perception, action, and cognition emerge from parallel computations among simple, interconnected unit...
By Lukas Braun, Erin Grant, Andrew M. Saxe