← Back to all news
arXiv Computation and Language August 27, 2026 By Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

Recurrence Meets Transformers for Universal Multimodal Retrieval

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • llms
  • rag
  • fine-tuning
  • multimodal
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Computer Vision
3d ago

Training-Free Pseudo-Fusion for Composed Image Retrieval with Diffusion Models and Multimodal Large Language Models

arXiv:2608.23102v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) is an emerging paradigm in content-based image retrieval that enables users to formulate compositional queries by combin...

By Fan Xu, Luis A. Leiva
llmsragdiffusionmultimodalbenchmarks
More like this →
arXiv AI
Aug 14

Generative Universal Multimodal Retrieval with Dual-role Identifiers

arXiv:2608. 12987v1 Announce Type: cross Abstract: Generative information retrieval (GIR) has emerged as a compelling alternative to the conventional index-retrieve-then-rank retrieval pipeline by training a generator to produce the identifiers of relevant items directly.

By Kaipeng Li, Haitao Yu, Xuanchen Zhou
ragdiffusionefficiencymultimodalbenchmarks
More like this →
arXiv AI
Jul 29

Multimodal Hybrid Retrieval-Augmented Generation for Scientific Document Understanding using Open-Source SLMs

arXiv:2607. 24799v1 Announce Type: cross Abstract: Large Language Models tend to hallucinate when answering domain-specific ques tions from scientific documents without prior fine-tuning.

By Alexandru-Andrei Sauc\u{a}, Ana-Luiza Rusnac
llmsragfine-tuningmultimodalbenchmarks
More like this →
arXiv AI
Jun 4

Overview of the EReL@MIR 2025 Multimodal Document Retrieval Challenge (Track 1)

arXiv:2606. 04240v1 Announce Type: cross Abstract: Retrieval over visually-rich documents, pages that interleave text with figures, tables, and charts, is essential for multimodal retrieval-augmented generation, yet most retrievers still discard the visual channel.

By Jingbiao Mei
llmsragfine-tuningmultimodal
More like this →
arXiv AI
Jun 16

MMLongEmbed: Benchmarking Multimodal Embedding Models in Long-Context Scenarios

arXiv:2606. 14747v1 Announce Type: cross Abstract: Recent advancements have significantly expanded the theoretical context windows of Multimodal Embedding Models (MEMs).

By Haitian Wang, Ruoxi Sun, Quantong Qiu, Juntao Li, Junhui Li, Hua Chen, Jinxiong Chang, Min Zhang
ragmultimodalbenchmarks
More like this →
arXiv AI
Jun 10

One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA

arXiv:2606. 10572v1 Announce Type: new Abstract: External memory effectively grounds large language models (LLMs) and vision-language models (VLMs)-based question answering (QA) in relevant multimodal evidence.

By Zhi Zheng, Ziqiao Meng, Hao Luan, Wei Liu, Wee Sun Lee
llmsragnlpefficiencymultimodalbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea