Multimodal Representation Alignment for Cross-modal Information Retrieval
arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.
arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.
The paper investigates how the modality gap— the separation between image and text representations in contrastive vision‑language models—affects different downstream tasks. By showing that a single dominant direction accounts for most of the image‑text mean separation, the authors explain why reducing or removing this gap can improve zero‑shot classification, degrade retrieval, or restore performance depending on the task. The study provides a geometric framework that clarifies when and why gap interventions should be applied in vision‑language systems.
arXiv:2607. 25266v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible.
arXiv:2609.28222v1 Announce Type: new Abstract: Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to...
arXiv:2608.29313v1 Announce Type: cross Abstract: CLIP-like vision-language models (VLMs) trained with contrastive objectives learn strong global image-text representations, but their Euclidean embed...
The paper introduces SCALAR, a query‑conditioned spherical centroid aggregator that assigns relevance‑based weights to each available modality before computing a spherical centroid. SCALAR supports arbitrary modality subsets, is trained with rank‑8 LoRA adapters on masked views, and achieves positive aggregation gains on four of five benchmarks, outperforming prior symmetric aggregators. With only 4.8 million trainable parameters, SCALAR attains the highest text‑to‑video R@1 on three benchmarks and surpasses the released GRAM checkpoint under test‑time modality dropout by 3.2 to 10.9 R@1.
arXiv:2606. 15134v1 Announce Type: cross Abstract: Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a scalar that uniformly pushes the embedding apart or pulls it together, as if every visual attribute either differed or matched.
arXiv:2609.15152v1 Announce Type: cross Abstract: Multimodal embedding models encode heterogeneous inputs into a shared embedding space, enabling efficient similarity computation across modalities an...
arXiv:2609.37225v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods eith...
Upgrading embedding models typically requires expensive database re-indexing, as new query embeddings are incompatible with existing database embeddings. While Backward Compatible Training (BCT) mitig...
arXiv:2606. 29464v1 Announce Type: cross Abstract: Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language models under strict data and compute budgets.
arXiv:2509. 07295v4 Announce Type: replace-cross Abstract: Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture.