arXiv Computer Vision

Hypergraph-Regularized Gramian Volumes for Multimodal Retrieval

arXiv Computer Vision
Sep 15

Query-Conditioned Spherical Centroid Aggregation for Multimodal Retrieval

The paper introduces SCALAR, a query‑conditioned spherical centroid aggregator that assigns relevance‑based weights to each available modality before computing a spherical centroid. SCALAR supports arbitrary modality subsets, is trained with rank‑8 LoRA adapters on masked views, and achieves positive aggregation gains on four of five benchmarks, outperforming prior symmetric aggregators. With only 4.8 million trainable parameters, SCALAR attains the highest text‑to‑video R@1 on three benchmarks and surpasses the released GRAM checkpoint under test‑time modality dropout by 3.2 to 10.9 R@1.

By Ambuj Mehrish, Anindya Nag, Sebastiano Vascon
arXiv Machine Learning
Jun 16

MVEB: Massive Video Embedding Benchmark

arXiv:2606. 14958v1 Announce Type: cross Abstract: We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classification, retrieval, and video-centric question answering.

By Adnan El Assadi, Roman Solomatin, Isaac Chung, Chenghao Xiao, Deep Shah, Manan Dey, Shriya Sudhakar, Zacharie Bugaud, Wissam Siblini, Ayush Sunil Munot, Yashwanth Devavarapu, Rakshitha Ireddi, Michelle Yang, M\'arton Kardos, Niklas Muennighoff, Kenneth Enevoldsen
arXiv AI
Aug 18

Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement

arXiv:2608. 16628v1 Announce Type: new Abstract: Modern Multimodal Retrieval-Augmented Generation (M-RAG) systems are fundamentally limited by the binary connectivity paradigm of traditional simple graphs, which fails to capture the intricate, high-order correlations among heterogeneous entities, such as the N-ary relationships between a visual chart, its scattered textual descriptions, and underlying numerical data.

By Shenao Chen, Yidan Xu, Xiangmin Han, Rundong Xue, Duanpo Wu, Yuhan Gao, Chenggang Yan, Yue Gao
arXiv AI
Aug 20

Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval

The paper introduces HN-CLIP, a new objective for dense-caption retrieval that adapts similarity margins per negative example using the text encoder’s own geometry. By adding a detached caption‑similarity matrix to the negative logits, HN‑CLIP addresses the issue of near‑duplicate captions that cause premature loss saturation in InfoNCE training. Experiments on four benchmarks show that HN‑CLIP outperforms leading methods by 2.4–4.3 R@1, trains 2.4× faster than GOAL and 5.4× faster than StructXLIP, and achieves the best full‑data baseline with only 20% of the training data.

By Haoyue Liu, Ye Chen, Zhichao Wang, Xiaoying Tang
arXiv Computer Vision
Sep 3

MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval?

MARS introduces a multi‑layer, multi‑slot embedding framework for text‑video retrieval that constructs adaptive representation slots by combining hidden states from different decoder layers. By comparing corresponding text and video slots and aggregating their similarities, MARS captures fine‑grained cues that single‑token embeddings miss. A hard‑negative‑aware slot specialization objective further encourages slots to focus on discriminative matching cues, leading to state‑of‑the‑art results on four benchmarks.

By Uicheol Jung, Juyoung Hong, Geuntaek Lim, Yukyung Choi