← Back to all news
arXiv Computer Vision September 15, 2026 By Amirul Rahman, Aisha Karim, Kenji Nakamura, Yi-Fan Ng

Capability-Routed Visual Retrieval and Evidence Threading for Long-Context Document Question Answering

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • rag
  • computer-vision
  • nlp

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Aug 18

What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering

arXiv:2608. 14841v1 Announce Type: new Abstract: Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text, tables, charts, and figures typically follows retrieve-then-read pipelines.

By Guanchen Wu, Jiayuan Ding, Subhabrata Mukherjee, Carl Yang
llmsnlpmultimodal
More like this →
arXiv AI
Aug 21

Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

arXiv:2608. 19739v1 Announce Type: cross Abstract: Multimodal LLMs can see a document, but they often can't read it reliably.

By Alin-Ionut Popa
llmsagentsnlpmultimodal
More like this →
arXiv AI
4d ago

Iterative Multimodal Retrieval-Augmented Generation for Medical Question Answering

arXiv:2604.27724v2 Announce Type: replace Abstract: Medical retrieval-augmented generation (RAG) systems typically operate on text chunks extracted from biomedical literature, discarding the rich vis...

By Xupeng Chen, Binbin Shi, Chenqian Le, Jiaqi Zhang, Kewen Wang, Ran Gong, Jinhan Zhang, Chihang Wang
llmsragcomputer-visionnlpmultimodalbenchmarks
More like this →
Hugging Face Trending Papers
Aug 20

Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual cues, and topological elements still trip them up under direct visual inference, even when the pag...

llmsagentsnlpmultimodal
More like this →
arXiv Machine Learning
Jun 8

FLOWREADER: Min-Cost Flow Optimization for Multi-Modal Long Document Q&A

arXiv:2606. 07235v1 Announce Type: cross Abstract: Long, multimodal documents force retrieval-augmented systems to assemble answers from evidence fragmented across text, tables, and slides broken across cells in a long table, spread over multiple slides, or split between a figure and its discussion.

By Ambuj Mehrish, Sebatiano Vascon
ragmultimodal
More like this →
arXiv AI
Jun 16

MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA

arXiv:2606. 15906v1 Announce Type: cross Abstract: Long-document multimodal question answering requires a system to locate sparse evidence in long PDFs and integrate clues from text, tables, images, charts, and complex layouts.

By Yilong Zuo, Xunkai Li, Jing Yuan, Qiangqiang Dai, Hongchao Qin, Ronghua Li
ragagentsnlpmultimodal
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea