Language Models Represent and Transform Concepts with Shared Geometry
How concepts are represented in neural networks is a fundamental question in machine learning. The dominant view treats concept representations as stationary geometric objects.
arXiv:2607. 04525v1 Announce Type: cross Abstract: How concepts are represented in neural networks is a fundamental question in machine learning.
How concepts are represented in neural networks is a fundamental question in machine learning. The dominant view treats concept representations as stationary geometric objects.
arXiv:2602. 15029v3 Announce Type: replace Abstract: The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes can be decoded using a linear probe.
arXiv:2607. 10578v1 Announce Type: new Abstract: Existing hypotheses represent a concept in an LLM as a single point, a linear direction, or a Gaussian cluster, yet it remains unclear how and why such structures emerge.
The paper studies how transformer representations evolve across layers by examining the intrinsic dimensionality (ID) of token embeddings and their neighborhood structures. It finds that closed‑class tokens expand and collapse earlier than open‑class tokens, and that these changes are linked to shifts in local geometry. The authors compare encoder and decoder models, showing distinct layer‑wise behaviors, and demonstrate that geometric features alone can predict a token’s part‑of‑speech and reveal how semantic content changes across layers.
The paper extends mechanistic interpretability of large language models by modeling concepts as low‑dimensional non‑linear manifolds rather than linear subspaces. It introduces a concept‑based alignment (CBA) score to compare these manifolds across layers and models, revealing block structures in intermediate layers, a shift from syntax‑dominated to mixed syntactic‑semantic concepts, and training‑dependent multilingual sharing. The study also shows that alignment patterns differ across model families and training stages, with adjacent stages aligning more closely than distant ones.
The paper investigates how large language models (LLMs) share a common Fisher‑Rao geometry in their next‑token probability distributions, revealing that behaviour largely determines this geometry while activation geometry depends on coordinate choices. Across transformer, state‑space, and recurrent architectures, output geometries align more closely than activation geometries, and this shared structure facilitates semantic‑category transfer and improves agreement with human word choices as models scale and train. The study further demonstrates that geometry can guide minimum‑disturbance interventions, enabling reusable control that preserves behaviour better than Euclidean methods and enhances steering, editing, attribution, dictionary learning, and fine‑tuning.
arXiv:2603. 06592v2 Announce Type: replace-cross Abstract: Contemporary studies in mechanistic interpretability have uncovered many puzzling phenomena in the neural information processing of Transformer-based language models, such as induction heads, function vectors, and the Hydra effect.
arXiv:2609.24209v1 Announce Type: new Abstract: The Platonic Representation Hypothesis (PRH) claims that independently trained models converge on a shared statistical model of reality, yet recent wor...
arXiv:2609.24821v1 Announce Type: new Abstract: The Linear Representation Hypothesis associates high-level concepts with directions in language models, but it remains unclear how these concept-relate...
The Platonic Representation Hypothesis (PRH) claims that independently trained models converge on a shared statistical model of reality, yet recent work finds only weak pointwise similarity between mo...
arXiv:2608.28980v1 Announce Type: cross Abstract: Can the specialized architectures that machine learning has traditionally built for structured data be replaced by language-based models? This questi...
arXiv:2602. 24264v2 Announce Type: replace-cross Abstract: Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems.