arXiv:2606. 07235v1 Announce Type: cross Abstract: Long, multimodal documents force retrieval-augmented systems to assemble answers from evidence fragmented across text, tables, and slides broken across cells in a long table, spread over multiple slides, or split between a figure and its discussion.
By Ambuj Mehrish, Sebatiano Vascon
arXiv:2607. 05438v1 Announce Type: cross Abstract: Multimodal retrieval-augmented generation (RAG) grounds a generator in evidence drawn from heterogeneous modalities -- text, tables, and images.
By Xue Li, Yiming Gai
The paper demonstrates that a single-pass multimodal model struggles to produce faithful, comprehensive reviews of long recordings or documents, often omitting a third of the content and embellishing the rest. By splitting the task into two passes—first transcribing the source and then reviewing the transcript—the authors show improved faithfulness and coverage across a diverse set of 21 sources. The benefit is most pronounced for longer or weaker baseline cases, while the approach introduces new failure modes such as space constraints and memory confabulation.
By Bojie Li, Noah Shi
arXiv:2607. 25422v1 Announce Type: new Abstract: Knowledge-intensive multimodal question answering (KI-MMQA) sits at the intersection of three expensive primitives: long visual token sequences, dense retrieval over large external corpora, and full cross-modal fusion.
By Noor Islam S. Mohammad, Ulu\u{g} Bayaz{\i}t
arXiv:2608. 19739v1 Announce Type: cross Abstract: Multimodal LLMs can see a document, but they often can't read it reliably.
By Alin-Ionut Popa
arXiv:2608.05592v2 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) have made strong progress in video understanding, yet long videos remain difficult: the visual token budge...
By Ziling Huang, Shin'ichi Satoh