arXiv Machine Learning By Liam Storan, Andreas Tolias, Nina Miolane

Group-Invariant Statistics Determine Embedding Geometry: Harmonic Analysis of Representations from Bach to the Night Sky

Read the original on arXiv Machine Learning →

The paper shows that the geometric patterns seen in language model embeddings—such as circles for months and saddle-shaped manifolds—arise from group-invariant statistics in word co‑occurrence data. By extending previous work on translation symmetry to arbitrary finite, compact, and homogeneous groups, the authors prove that embeddings correspond to matrix elements of the irreducible representations of the symmetry group. They validate this theory experimentally with the cyclic group <Z_{12}> for months, a dihedral group for musical chords, and a spherical‑harmonic embedding for celestial objects.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 1

Symmetry in language statistics shapes the geometry of model representations

arXiv:2602. 15029v3 Announce Type: replace Abstract: The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes can be decoded using a linear probe.

By Dhruva Karkada, Daniel J. Korchinski, Andres Nava, Matthieu Wyart, Yasaman Bahri
arXiv AI
4d ago

Causal and Interpretable Structures in LLM Compositional Tasks

The paper investigates how large language models encode and use relational information among tokens across transformer layers. By analyzing activations from prompts that require inferring relationships among three cyclic tokens (months, hours, weekdays, musical notes), the authors find a consistent layerwise progression: intermediate layers capture pairwise relationships, while later layers encode the full three‑token relationship to predict the next token. They also identify geometrically structured token relationships that do not influence prediction, and show that constraining models to use only causally relevant joint representations improves next‑token accuracy.

By Gurbir Arora, Toni J. B. Liu, Jiajun Bao, Rapha\"el Sarfati, Christopher J. Earls