The standard way to compare two text embeddings is cosine similarity. Scattered studies report that a different metric does better, but never pin down the geometric condition that decides when, or why.
arXiv:2602. 19393v2 Announce Type: replace Abstract: Steck, Ekanadham, and Kallus [arXiv:2403.
By Taha Bouhsine
arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.
By Fan Xu, Luis A. Leiva
arXiv:2606. 30625v1 Announce Type: cross Abstract: Contrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosine similarity, effectively ignoring embedding magnitudes.
By Ziwei Su, Junyu Ren, Victor Veitch
arXiv:2606. 15054v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) detect features via inner product, so a feature's activation scales with both its directional alignment and the input's norm.
By Silen Naihin, Lev Stambler
arXiv:2606. 28330v1 Announce Type: cross Abstract: Embedding-based retrieval systems rely on the assumption that geometric proximity in highdimensional representation spaces reflects semantic relevance.
By Ernesto Lopez Fune (DE)