← Back to all news
arXiv Computer Vision August 25, 2026 By Qiyou Liu, Yong Zhang, Jianjie Luo, Zhenguo Yang, Yi Yu

Boosting Knowledge-based Visual Question Answering with Structured Context Reasoning

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • llms
  • nlp
  • efficiency
  • multimodal
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
1d ago

SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering

arXiv:2608.21796v1 Announce Type: cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) aims to answer queries that necessitate reasoning over external knowledge sources beyond the visua...

By Long Shu, Shuochen Liu, Wei Chen, Junda Lin, Zhi Zheng, Huijun Hou, Tong Xu
llmsnlpreinforcement-learningmultimodalbenchmarkssafety
More like this →
arXiv AI
1d ago

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

arXiv:2608.22429v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However, this relian...

By Changjiang Jiang, Qiannian Zhao, Lei Xin, Jinxiang Xie, Preslav Nakov, Zhuohan Xie
llmsreinforcement-learningfine-tuningefficiencymultimodal
More like this →
arXiv Computer Vision
1d ago

IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning

arXiv:2608.23330v1 Announce Type: new Abstract: Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human action...

By Jiapeng Li, Ping Wei, Wenjuan Han, Song-Chun Zhu, Lifeng Fan
llmsagentsnlpmultimodalbenchmarks
More like this →
arXiv AI
Jul 21

MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs

arXiv:2507. 20804v3 Announce Type: replace Abstract: Large Language Models (LLMs) suffer from hallucinations due to their static parametric knowledge.

By Xueyao Wan, Hang Yu
llmsragmultimodalbenchmarkssafety
More like this →
arXiv AI
Aug 7

Look Twice: Training-Free Evidence Highlighting for Knowledge-based Visual Question Answering

arXiv:2604. 01280v2 Announce Type: replace-cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires Multimodal Large Language Models (MLLMs) to identify and combine fine-grained visual cues with retrieved textual evidence.

By Marco Morini, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
llmsnlpmultimodalbenchmarks
More like this →
arXiv AI
Jun 2

Multimodal Function Vectors for Visual Relations

arXiv:2510. 02528v2 Announce Type: replace Abstract: Large Multimodal Models (LMMs) demonstrate impressive in-context learning abilities from few multimodal demonstrations, yet the internal mechanisms supporting such task learning remain opaque.

By Shuhao Fu, Esther Goldberg, Ying Nian Wu, Hongjing Lu
llmsfine-tuningmultimodal
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea