arXiv AI

CMDR: Contextual Multimodal Document Retrieval

arXiv:2607. 05927v1 Announce Type: cross Abstract: Multimodal document retrieval aims to retrieve relevant pages while preserving both textual and visual content from the original document.

arXiv Computation and Language
Sep 23

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding

CausalEmbed is an auto‑regressive method for generating compact multi‑vector embeddings in visual document retrieval. By using iterative margin loss during contrastive training, it reduces the number of visual tokens needed by 30‑155× while keeping performance competitive across different backbones and benchmarks. The approach offers efficient training, scalable test‑time performance, and a flexible scaling strategy for multi‑vector representations.

By Jiahao Huo, Yu Huang, Yibo Yan, Ye Pan, Kening Zheng, Wei-Chieh Huang, Yi Cao, Mingdong Ou, Philip S. Yu, Xuming Hu
arXiv Computation and Language
Sep 18

Less Is More: Graph-free Multimodal RAG via Multi-signal Late Fusion

The paper introduces TrioRAG, a graph-free multimodal retrieval-augmented generation framework that combines evidence from the question, an anchor image, and a VLM-enhanced query via late fusion. It also presents AutoQA, a benchmark featuring noisy web-sourced images that require reasoning across manuals. TrioRAG outperforms graph-based systems on three benchmarks while cutting costs and speeding up inference by 1.6–2.3×.

By Tithi Rakshit, Hongkuan Zhou, Lavdim Halilaj, Yuqicheng Zhu
arXiv AI
Aug 18

Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement

arXiv:2608. 16628v1 Announce Type: new Abstract: Modern Multimodal Retrieval-Augmented Generation (M-RAG) systems are fundamentally limited by the binary connectivity paradigm of traditional simple graphs, which fails to capture the intricate, high-order correlations among heterogeneous entities, such as the N-ary relationships between a visual chart, its scattered textual descriptions, and underlying numerical data.

By Shenao Chen, Yidan Xu, Xiangmin Han, Rundong Xue, Duanpo Wu, Yuhan Gao, Chenggang Yan, Yue Gao
arXiv Machine Learning
Jun 3

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning

arXiv:2603. 01471v3 Announce Type: replace-cross Abstract: Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across diverse tasks such as retrieval and classification.

By Jiahan Chen, Da Li, Hengran Zhang, Yinqiong Cai, Lixin Su, Jiafeng Guo, Daiting Shi, Dawei Yin, Keping Bi
arXiv Machine Learning
Jun 2

Reconstructing Content via Collaborative Attention to Improve Multimodal Embedding Quality

arXiv:2603. 01471v2 Announce Type: replace-cross Abstract: Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across diverse tasks such as retrieval and classification.

By Jiahan Chen, Da Li, Hengran Zhang, Yinqiong Cai, Lixin Su, Jiafeng Guo, Daiting Shi, Dawei Yin, Keping Bi
arXiv Computation and Language
Aug 27

Recurrence Meets Transformers for Universal Multimodal Retrieval

The paper introduces ReT-2, a unified retrieval model that handles multimodal queries containing both images and text and searches across multimodal document collections. It employs a recurrent Transformer architecture with LSTM-inspired gating to integrate information across layers and modalities, capturing fine-grained visual and textual details. Evaluations on M2KR and M-BEIR benchmarks show state‑of‑the‑art performance, faster inference, and lower memory usage, and the model also boosts downstream tasks in retrieval‑augmented generation pipelines.

By Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
arXiv Computer Vision
Sep 1

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings

arXiv:2604.22280v4 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have emerged as a promising foundation for universal multimodal embeddings. Recent studies have shown that...

By Peixi Wu, Ke Mei, Feipeng Ma, Bosong Chai, Zhibin Lan, Chenxi Zhao, Shannan Yan, Jie Chen, Zhangchi Hu, Yansong Peng, Bo Lin, Junjie Zhou, Dacheng Yin, Tianyi Wang, Fengyun Rao, Jing Lyu, Hebei Li, Xiaoyan Sun