The standard way to compare two text embeddings is cosine similarity. Scattered studies report that a different metric does better, but never pin down the geometric condition that decides when, or why.
arXiv:2602. 19393v2 Announce Type: replace Abstract: Steck, Ekanadham, and Kallus [arXiv:2403.
By Taha Bouhsine
arXiv:2609.15152v1 Announce Type: cross
Abstract: Multimodal embedding models encode heterogeneous inputs into a shared embedding space, enabling efficient similarity computation across modalities an...
By Yanping Li, Wei Zhou, Yawen Liu, Yibo Wang, Ke Zhu, Guangda Huzhang, Qing-Guo Chen, Zhao Xu, Jun Zhang, Wei Wei
arXiv:2609.39836v1 Announce Type: new
Abstract: Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natura...
By Simone Ricci, Niccol\`o Biondi, Federico Pernici
arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.
By Fan Xu, Luis A. Leiva
arXiv:2608.30263v1 Announce Type: cross
Abstract: Large vision-language models (LVLMs) incur substantial inference costs due to their long and highly redundant visual-token sequences. Diversity-based...
By Shunjie Wen, Jaeyeon Lee, Dong-Wan Choi