← Back to all news
Hugging Face Trending Papers August 20, 2026

Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

Read the original on Hugging Face Trending Papers →

The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.

  • llms
  • agents
  • nlp
  • multimodal

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Aug 21

Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

arXiv:2608. 19739v1 Announce Type: cross Abstract: Multimodal LLMs can see a document, but they often can't read it reliably.

By Alin-Ionut Popa
llmsagentsnlpmultimodal
More like this →
arXiv AI
Jul 20

Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models

arXiv:2607. 15565v1 Announce Type: cross Abstract: Where should the question go in a vision-language model (VLM) prompt: before the image or after it?

By Rakshanda Hassan Abhinandan, John Galeotti, Deva Ramanan, Gautam Rajendrakumar Gare
llmsnlpfine-tuningmultimodalbenchmarks
More like this →
arXiv AI
Jun 16

Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering

arXiv:2606. 16494v1 Announce Type: cross Abstract: Knowledge-based visual question answering (KB-VQA) lets vision-language systems answer questions that exceed their parametric knowledge by conditioning a reader on passages retrieved from a Wikipedia-scale knowledge base.

By Jieyuan Liu, Jianyang Gu, Shijie Chen, Jefferson Chen, Zhen Wang
llmsragnlpmultimodalbenchmarkssafety
More like this →
arXiv AI
Aug 11

Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution

arXiv:2608. 07943v1 Announce Type: new Abstract: Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window.

By Lewei Xu, Yihao Ding, Zihan Xu, Daniel Yitian Su, Daochang Liu, Siwen Luo, Yifan Peng, Wei Liu
llms
More like this →
arXiv AI
Jul 8

Modality Relevance is not Modality Utility: Post-hoc Selective Modality Escalation for Cost-Aware Multimodal RAG

arXiv:2607. 05438v1 Announce Type: cross Abstract: Multimodal retrieval-augmented generation (RAG) grounds a generator in evidence drawn from heterogeneous modalities -- text, tables, and images.

By Xue Li, Yiming Gai
llmsragmultimodal
More like this →
arXiv Computer Vision
Aug 26

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning

arXiv:2606.18974v3 Announce Type: replace Abstract: Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly a...

By Pengyu Li, Zhitao Gao, Lingling Zhang, Muye Huang, Yuanming Li, Fangzhi Xu, Jun Liu
diffusionefficiencymultimodalbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea