arXiv Machine Learning By Dhruva Karkada, Daniel J. Korchinski, Andres Nava, Matthieu Wyart, Yasaman Bahri

Symmetry in language statistics shapes the geometry of model representations

Read the original on arXiv Machine Learning →

arXiv:2602. 15029v3 Announce Type: replace Abstract: The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes can be decoded using a linear probe.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

Group-Invariant Statistics Determine Embedding Geometry: Harmonic Analysis of Representations from Bach to the Night Sky

The paper shows that the geometric patterns seen in language model embeddings—such as circles for months and saddle-shaped manifolds—arise from group-invariant statistics in word co‑occurrence data. By extending previous work on translation symmetry to arbitrary finite, compact, and homogeneous groups, the authors prove that embeddings correspond to matrix elements of the irreducible representations of the symmetry group. They validate this theory experimentally with the cyclic group <Z_{12}> for months, a dihedral group for musical chords, and a spherical‑harmonic embedding for celestial objects.

By Liam Storan, Andreas Tolias, Nina Miolane