arXiv AI By Archan Dutta, Vyanktesh Kanungo

Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval

Read the original on arXiv AI →

arXiv:2608. 11343v1 Announce Type: new Abstract: Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 3

MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval?

MARS introduces a multi‑layer, multi‑slot embedding framework for text‑video retrieval that constructs adaptive representation slots by combining hidden states from different decoder layers. By comparing corresponding text and video slots and aggregating their similarities, MARS captures fine‑grained cues that single‑token embeddings miss. A hard‑negative‑aware slot specialization objective further encourages slots to focus on discriminative matching cues, leading to state‑of‑the‑art results on four benchmarks.

By Uicheol Jung, Juyoung Hong, Geuntaek Lim, Yukyung Choi
arXiv Computation and Language
Aug 27

Recurrence Meets Transformers for Universal Multimodal Retrieval

The paper introduces ReT-2, a unified retrieval model that handles multimodal queries containing both images and text and searches across multimodal document collections. It employs a recurrent Transformer architecture with LSTM-inspired gating to integrate information across layers and modalities, capturing fine-grained visual and textual details. Evaluations on M2KR and M-BEIR benchmarks show state‑of‑the‑art performance, faster inference, and lower memory usage, and the model also boosts downstream tasks in retrieval‑augmented generation pipelines.

By Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara