Hugging Face Trending Papers

Geometry of Ordinal Representations in Language Models

Recent work showed that language models represent character counts on curved 1D manifolds, with attention heads performing geometric transformations to enable computation. We test whether this generalizes across four ordinal tasks (bracket depth, indentation, table position, numeric magnitude) in Gemma-2-2B, Gemma-2-9B, and Qwen3-4B.

arXiv Computation and Language
Sep 7

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

The paper investigates how large language models (LLMs) organize reasoning operations—such as problem formulation, goal decomposition, and deduction—within their hidden representation spaces. It shows that these operations are separable in held‑out representations, with peak separability in middle layers, and that token‑wise alignment of operations becomes more distributed across spans as layers deepen. Attention‑masking experiments reveal that representations aligned to operations at chunk onsets depend on prior reasoning context, indicating a geometric correspondence between linguistic reasoning expressions and internal model structure.

By Seogyeong Jeong, Jaehui Hwang, Dongyoon Han, Geonmo Gu, Alice Oh, Taekyung Kim
arXiv AI
Aug 19

Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

The paper audits frozen decoder‑only large language models (LLMs) on geometric reasoning tasks using parametric CAD constraints. It probes hidden states for linear decodability, forced‑choice generation, activation‑level influence, and behavioral steerability, finding that pretraining improves decoding of local geometric relations but not sketch‑level DOF status. The study shows that decodable information is not always actionable: generation often fails to express it, and steering interventions do not reliably control outputs, revealing divergences among decodability, generation, activation influence, and steerability.

By Man Liang, Xinzhao Cheng, Faizan Wajid
arXiv AI
Sep 3

Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders

The paper proposes an encoder Transformer that explicitly separates semantic, absolute positional (AP), and relative positional (RP) information, restricting the masked‑language‑modeling objective to the semantic stream. This disentanglement reveals that the AP subspace collapses into a low‑frequency two‑dimensional manifold reflecting document structure, that attention heads specialize into structure‑ and semantic‑oriented groups with RP supporting only the latter, and that standard positional encodings fail to robustly encode macroscopic structure. The approach preserves positional encoding and improves performance on 49 out of 65 linguistic phenomena in the Flash‑Holmes probing benchmark.

By Pierre-Antoine Lequeu, Camille Barboule, Benjamin Piwowarski
arXiv AI
4d ago

Causal and Interpretable Structures in LLM Compositional Tasks

The paper investigates how large language models encode and use relational information among tokens across transformer layers. By analyzing activations from prompts that require inferring relationships among three cyclic tokens (months, hours, weekdays, musical notes), the authors find a consistent layerwise progression: intermediate layers capture pairwise relationships, while later layers encode the full three‑token relationship to predict the next token. They also identify geometrically structured token relationships that do not influence prediction, and show that constraining models to use only causally relevant joint representations improves next‑token accuracy.

By Gurbir Arora, Toni J. B. Liu, Jiajun Bao, Rapha\"el Sarfati, Christopher J. Earls
arXiv Machine Learning
Jul 1

Symmetry in language statistics shapes the geometry of model representations

arXiv:2602. 15029v3 Announce Type: replace Abstract: The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes can be decoded using a linear probe.

By Dhruva Karkada, Daniel J. Korchinski, Andres Nava, Matthieu Wyart, Yasaman Bahri