arXiv AI

MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

arXiv:2608. 09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion.

arXiv AI
Sep 12

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Sci‑MMR is a new benchmark for multi‑step evidence‑grounded scientific reasoning in multimodal agents, featuring 235 multi‑hop tasks across four disciplines and an average of nine figure panels per task. It evaluates not just final answer accuracy but also the recovery of structured evidence from scientific claims, citations, visual data, and supporting regions. Experiments on eight state‑of‑the‑art models show a gap of over 20 points between answer accuracy and complete evidence recovery, highlighting significant challenges in evidence acquisition and integration.

By Jiaqiang Li, Yajie Yang, Zhiheng Xi, Jiadong Chen, Enyu Zhou, Senjie Jin, Yang Nan, Jiazheng Zhang, Han Wang, Yanxin Li, Dingwei Zhu, Bicheng Deng, Yuhui Wang, Xiang Zheng, Qi Zhang, Lei Bai, Xingjun Ma, Tao Gui
arXiv AI
Aug 26

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes

The paper introduces CAIT, a benchmark of 400 synthetic scenes featuring counter‑intuitive actions that challenge multimodal large language models (MLLMs). Human participants and proprietary models like Claude and Gemini perform well, but standard open‑source instruction‑tuned MLLMs fail, largely due to a strong language prior that overrides contradictory visual evidence. The study shows that Chain‑of‑Thought reasoning can help but introduces new issues, while targeted fine‑tuning and structured prompting can reduce reliance on language priors and improve visual grounding.

By Chen Ling, Tongwei Zhang, Hanqian Li, Nai Ding
arXiv Computer Vision
Sep 16

MechReason: Benchmarking Multi-Image Multi-Hop Reasoning in Mechanical Engineering

MechReason is a new benchmark for multi-image, multi-hop reasoning in mechanical engineering, featuring 12,000 question-answer pairs with explicit reasoning chains and 21,000 visual items across nine evidence types. It covers eight task types—explanation, prediction, design, and diagnosis—within four reasoning dimensions, and is constructed through a four-stage pipeline that extracts engineering claims, generates verification-masked questions, and validates multimodal quality. Even state‑of‑the‑art models score only 62.89% on this dataset, highlighting its difficulty.

By Tengyue Wang, Kang An, Chenxu Du, Zhongyu Yang, Yuanchi Zhu, Xinqi Yang, Hebao Zhu, Ziliang Wang, FaQiang Qian, Yunli Yang, Qibing Ren
arXiv AI
Sep 12

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Mr.LHDR is a new benchmark designed to evaluate deep research agents on long‑horizon, multimodal tasks. It presents questions built from hidden Node‑Relation graphs that require an average of 12.1 intermediate conclusions and a mean dependency depth of 10.4 before arriving at a single verifiable answer. The benchmark tests both final answers and the correctness of intermediate conclusions, using metrics such as Overall Accuracy, Strict Accuracy, Checklist Score, and Dependency‑Aware Checklist Score.

By Minghao Guo, Meng Cao, Sui Zhao, Siyu Ning, Xin Wang, Haoze Zhao, Jiaxuan Yang, Haihong Hao, Mingfei Han, Shunlin Rong, Haijun Wu, Xiaodan Liang, Xiaojun Chang
arXiv AI
Jul 31

VistaHop: Benchmarking Long-Horizon Visual DeepSearch

arXiv:2606. 03273v2 Announce Type: replace-cross Abstract: Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evidence, and connecting fine-grained clues across multiple steps.

By Hang He, Chuhuai Yue, Chengqi Dong, Chengcheng Wan, Ting Su, Haiying Sun, Jiajun Chai, Xiaohan Wang, Guojun Yin
arXiv Computer Vision
Sep 22

Pay More Attention To Text In High-Resolution MLLMs

The paper introduces EviSpec, a training‑free compiler that generates complementary evidence specifications to improve high‑resolution multimodal large language models (MLLMs). By explicitly guiding visual search with structured evidence specifications, EviSpec achieves significant relative gains—up to 14.8% over random evidence—across five MLLMs and three benchmarks, and also sets new state‑of‑the‑art results on VQA and hallucination‑focused tasks.

By Zhongkuan Mao, Wenzhuo Zhao, Xianjie Liu, Yidong Wang, Zhao Gao, Ronghao Xian, Yao Jiang, Yi Zhang, Liangjian Wen, Keren Fu
arXiv Computation and Language
Sep 23

ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains

The ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains introduced a new Visual Question Answering benchmark that tests reasoning over documents from eight distinct domains such as business reports, scientific papers, and engineering drawings. Twenty valid submissions from eight teams were evaluated, featuring approaches ranging from zero‑shot vision‑language models to multi‑agent ensembles and fine‑tuned multimodal systems. Results indicate that the most effective systems employ structured evidence extraction, retrieval, verification, and orchestration across multiple components rather than single‑pass prompting.

By Artemis Llabr\'es, Marc Serra Ortega, Tom\`as Ockier, Samuel Ortega Cuadra, Amritpal Singh, Christos Georgakilas, Andrey Barsky, Ernest Valveny, Dimosthenis Karatzas
arXiv Computation and Language
Sep 1

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

The paper surveys Multimodal Code Intelligence, focusing on tasks where code is generated, edited, refined, or reasoned about under visually grounded inputs such as screenshots, charts, and videos. It categorizes the field by the role of code—rendered artifact, editable structure, intermediate reasoning trace, or executable tool interface—and organizes benchmarks into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Frameworks. The authors argue that reliable evaluation must include evidence of semantics and interaction beyond visual fidelity, and propose four verification-centered research directions to advance the field toward evidence-grounded executable systems.

By Xuanle Zhao, Qiushi Sun, Jingyu Xiao, Xuexin Liu, Haoyue Yang, Qiaosheng Chen, Xianzhen Luo, Jing Huang, Yufeng Zhong, Lei Chen, Shuai Fu, Zhenlin Wei, Jinhe Bi, Lei Jiang, Haibo Qiu, Siqi Yang, Peng Shi, Jian Hu, Zhixiong Zeng