arXiv Computation and Language
Sep 23

ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains

The ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains introduced a new Visual Question Answering benchmark that tests reasoning over documents from eight distinct domains such as business reports, scientific papers, and engineering drawings. Twenty valid submissions from eight teams were evaluated, featuring approaches ranging from zero‑shot vision‑language models to multi‑agent ensembles and fine‑tuned multimodal systems. Results indicate that the most effective systems employ structured evidence extraction, retrieval, verification, and orchestration across multiple components rather than single‑pass prompting.

By Artemis Llabr\'es, Marc Serra Ortega, Tom\`as Ockier, Samuel Ortega Cuadra, Amritpal Singh, Christos Georgakilas, Andrey Barsky, Ernest Valveny, Dimosthenis Karatzas
arXiv Computer Vision
Aug 31

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning

Doc‑CoB introduces a Chain‑of‑Boxes framework that enhances document understanding by progressively focusing on query‑relevant layout regions while preserving global context. It selects key layout boxes and then applies visual prompting for deeper analysis, supported by two new reasoning tasks and an automatic pipeline that generates 249k training samples with intermediate visual supervision. Experiments across seven benchmarks and four popular models demonstrate significant performance gains, underscoring the method’s effectiveness and broad applicability.

By Ye Mo, Kai Ye, Xianwei Mao, Zirui Shao, Gang Huang, Bo Zhang, Hangdi Xing, Kehan Chen, Huan Zhou, Zixu Yan, Jiajun Bu, Sheng Zhou
arXiv Computer Vision
Aug 27

Visual General Intelligence: A White Paper

The paper titled "Visual General Intelligence: A White Paper" reexamines intelligence from a vision-centered perspective, questioning whether visual experience and learning can lead to artificial general intelligence (AGI). It compares the success of language models like GPT, which transfer to unseen tasks via autoregressive modeling on large text corpora, with the potential of visual modalities such as images, videos, and geometry to develop similar capabilities. The authors aim to outline principles for computer vision in the AGI era, including input modalities, benchmarks, learning paradigms, and the interplay between vision and other modalities like language.

By Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian, Shangzhe Wu, Oishi Deb, Ryousuke Yamada, Christian Rupprecht, Jianyuan Wang, Kohsuke Ide, Koichi Namekata, Xianzheng Ma, Yiming Chen, Robert Geirhos, Aditi Raghunathan, Yuki M. Asano, Deva Ramanan, David Fouhey, Andrew J. Davison, Yilun Du, Jiajun Wu, Zhuang Liu