UOT-Gap: A Variational Principle for the Modality Gap in Vision-Language Models via Unbalanced Optimal Transport
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2602. 07026v3 Announce Type: replace-cross Abstract: Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions.
arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.
The paper introduces HN-CLIP, a new objective for dense-caption retrieval that adapts similarity margins per negative example using the text encoder’s own geometry. By adding a detached caption‑similarity matrix to the negative logits, HN‑CLIP addresses the issue of near‑duplicate captions that cause premature loss saturation in InfoNCE training. Experiments on four benchmarks show that HN‑CLIP outperforms leading methods by 2.4–4.3 R@1, trains 2.4× faster than GOAL and 5.4× faster than StructXLIP, and achieves the best full‑data baseline with only 20% of the training data.
arXiv:2606. 24716v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence.
arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.
arXiv:2609.05730v1 Announce Type: cross Abstract: Contrastive Language-Image Pretraining (CLIP) is a building block of many machine learning applications. Scaling laws have guided resource allocation...