MARS introduces a multi‑layer, multi‑slot embedding framework for text‑video retrieval that constructs adaptive representation slots by combining hidden states from different decoder layers. By comparing corresponding text and video slots and aggregating their similarities, MARS captures fine‑grained cues that single‑token embeddings miss. A hard‑negative‑aware slot specialization objective further encourages slots to focus on discriminative matching cues, leading to state‑of‑the‑art results on four benchmarks.
By Uicheol Jung, Juyoung Hong, Geuntaek Lim, Yukyung Choi
The paper introduces ReT-2, a unified retrieval model that handles multimodal queries containing both images and text and searches across multimodal document collections. It employs a recurrent Transformer architecture with LSTM-inspired gating to integrate information across layers and modalities, capturing fine-grained visual and textual details. Evaluations on M2KR and M-BEIR benchmarks show state‑of‑the‑art performance, faster inference, and lower memory usage, and the model also boosts downstream tasks in retrieval‑augmented generation pipelines.
By Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.
By Fan Xu, Luis A. Leiva
arXiv:2609.37225v1 Announce Type: cross
Abstract: Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods eith...
By Zijing Cai, Yuzhe Wang, Jingxian Zhu, Fengbin Zhu, Richang Hong
arXiv:2506. 03096v2 Announce Type: replace-cross Abstract: Contrastive language-image pre-training aligns features of text-image pairs in a common latent space via distinct encoders for each modality.
By Christian Schlarmann, Francesco Croce, Nicolas Flammarion, Matthias Hein
arXiv:2608.23102v1 Announce Type: new
Abstract: Composed Image Retrieval (CIR) is an emerging paradigm in content-based image retrieval that enables users to formulate compositional queries by combin...
By Fan Xu, Luis A. Leiva