arXiv Computer Vision

Spherical Interpolation for Backward-Compatible Multimodal Representations

arXiv AI
4d ago

One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models

The paper investigates how the modality gap— the separation between image and text representations in contrastive vision‑language models—affects different downstream tasks. By showing that a single dominant direction accounts for most of the image‑text mean separation, the authors explain why reducing or removing this gap can improve zero‑shot classification, degrade retrieval, or restore performance depending on the task. The study provides a geometric framework that clarifies when and why gap interventions should be applied in vision‑language systems.

By Aditya Sharma, Divya Saxena
arXiv Computer Vision
Sep 15

Query-Conditioned Spherical Centroid Aggregation for Multimodal Retrieval

The paper introduces SCALAR, a query‑conditioned spherical centroid aggregator that assigns relevance‑based weights to each available modality before computing a spherical centroid. SCALAR supports arbitrary modality subsets, is trained with rank‑8 LoRA adapters on masked views, and achieves positive aggregation gains on four of five benchmarks, outperforming prior symmetric aggregators. With only 4.8 million trainable parameters, SCALAR attains the highest text‑to‑video R@1 on three benchmarks and surpasses the released GRAM checkpoint under test‑time modality dropout by 3.2 to 10.9 R@1.

By Ambuj Mehrish, Anindya Nag, Sebastiano Vascon
arXiv AI
Jun 16

Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddings

arXiv:2606. 15134v1 Announce Type: cross Abstract: Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a scalar that uniformly pushes the embedding apart or pulls it together, as if every visual attribute either differed or matched.

By Shubhang Bhatnagar, Dheeraj Baiju, Narendra Ahuja