arXiv:2606. 30625v1 Announce Type: cross Abstract: Contrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosine similarity, effectively ignoring embedding magnitudes.
By Ziwei Su, Junyu Ren, Victor Veitch
arXiv:2609.05721v1 Announce Type: new
Abstract: Understanding whether language-model embeddings encode structured real-world information is important for both representation analysis and information...
By Esteban Feuerstein, Victoria Klimkowski, Juan Manuel Ortiz de Zarate, Federico Hern\'an Suaiter
arXiv:2606. 28330v1 Announce Type: cross Abstract: Embedding-based retrieval systems rely on the assumption that geometric proximity in highdimensional representation spaces reflects semantic relevance.
By Ernesto Lopez Fune (DE)
arXiv:2602. 15029v3 Announce Type: replace Abstract: The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes can be decoded using a linear probe.
By Dhruva Karkada, Daniel J. Korchinski, Andres Nava, Matthieu Wyart, Yasaman Bahri
arXiv:2608.28840v1 Announce Type: new
Abstract: Independently trained neural networks tend to encode the same data with similar latent geometries. These latent geometries are not directly compatible,...
By Cameron Ryan, Vivek Sivaraman Narayanaswamy, Kowshik Thopalli, Shusen Liu
We are excited to announce a new embedding model which is significantly more capable, cost effective, and simpler to use.
arXiv:2608. 06809v1 Announce Type: new Abstract: How can an analyst decide whether a nonlinear dimensionality reduction embedding can be trusted?
By Xinyu Zhang, Klaus Mueller
arXiv:2503.00612v2 Announce Type: replace-cross
Abstract: The basic problem of semantic compression is to minimize the length of a message while preserving its meaning. This differs from classical no...
By Tankut Can
The paper investigates the Platonic Representation Hypothesis, which posits that more capable models converge toward shared representations. By distinguishing relational structure (which samples are related) from metric geometry (quantitative relations like distances), the authors develop a controlled $2 imes2$ framework to evaluate both aspects at local and global scales. Their findings show that relational structure consistently converges across vision‑language and video‑text models, while metric geometry converges much more weakly, a pattern that persists even when using a Riemannian metric approximation.
By Junwon You, Mihyun Jang, Sangwoo Mo, Jae-Hun Jung
arXiv:2609.14917v1 Announce Type: cross
Abstract: We introduce document embedding geometry as a quantitative observable of conceptual reorganization and develop a counterfactual ablation framework fo...
By Dimitris Ntounis, Ariel Schwartzman, Chris Chafe, Thomas A. Ryckman
arXiv:2602. 14486v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis suggests that representations from neural networks are converging to a common statistical model of reality.
By Fabian Gr\"oger, Shuo Wen, Maria Brbi\'c
Retrieval-Augmented Generation systems rely on similarity scores to retrieve relevant content, yet scores are not directly comparable across embedding models due to differing geometric properties, complicating model migration and limiting threshold reuse. We study how similarity scores can be related by learning mappings between score distributions rather than embeddings.