arXiv AI

Emergent Ordinal Geometry in Transformers Trained on Local Comparisons

arXiv:2606. 01269v1 Announce Type: new Abstract: Transitive inference is the challenge of inferring that A < C from knowing only adjacent relations (A < B, B < C).

arXiv AI
4d ago

Causal and Interpretable Structures in LLM Compositional Tasks

The paper investigates how large language models encode and use relational information among tokens across transformer layers. By analyzing activations from prompts that require inferring relationships among three cyclic tokens (months, hours, weekdays, musical notes), the authors find a consistent layerwise progression: intermediate layers capture pairwise relationships, while later layers encode the full three‑token relationship to predict the next token. They also identify geometrically structured token relationships that do not influence prediction, and show that constraining models to use only causally relevant joint representations improves next‑token accuracy.

By Gurbir Arora, Toni J. B. Liu, Jiajun Bao, Rapha\"el Sarfati, Christopher J. Earls
arXiv AI
Jun 9

DiffoR: A Unified Continuous Generative Framework for Universal Ordinal Regression

arXiv:2606. 07599v1 Announce Type: cross Abstract: Ordinal Regression (OR) aims to predict target values with inherent order, underpinning critical applications across diverse domains, from recommender systems to computer vision.

By Hongxu Ma, Lin Wang, Chenghou Jin, Han Zhou, Jie Zhang, Xiaoyu Yang, Chunjie Chen, Jihong Guan, Shuigeng Zhou
arXiv Machine Learning
Sep 24

What Converges in the Platonic Representation Hypothesis? Structure over Geometry

The paper investigates the Platonic Representation Hypothesis, which posits that more capable models converge toward shared representations. By distinguishing relational structure (which samples are related) from metric geometry (quantitative relations like distances), the authors develop a controlled $2 imes2$ framework to evaluate both aspects at local and global scales. Their findings show that relational structure consistently converges across vision‑language and video‑text models, while metric geometry converges much more weakly, a pattern that persists even when using a Riemannian metric approximation.

By Junwon You, Mihyun Jang, Sangwoo Mo, Jae-Hun Jung