arXiv AI

HoloCount: A Holistic Visual Counting Benchmark for MLLMs

arXiv:2607. 06420v1 Announce Type: cross Abstract: Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning.

arXiv Computer Vision
Aug 26

Object Counting Across Modalities: Taxonomies, Benchmarks, Applications, and Open Challenges

The paper reviews the evolution of object‑counting techniques from class‑specific density regression to open‑vocabulary, foundation‑model‑backed counters that can handle visual and textual prompts. It highlights that current evaluation relies on a few saturated benchmarks, leading to models exploiting statistical regularities rather than true generalization. The authors propose a five‑axis taxonomy and audit the literature across domains such as microscopy, remote sensing, crowd counting, and agriculture, identifying six structural contradictions and outlining a roadmap for robust, multimodal evaluation protocols.

By Joana Konadu Owusu, Shivanand Venkanna Sheshappanavar
arXiv Computer Vision
Sep 3

Beyond Appearance: Can Multimodal Large Language Models Exploit Vertical Structure for Remote Sensing Natural Scene Understanding?

The paper introduces VertiCue-Bench, a diagnostic benchmark designed to test whether multimodal large language models (MLLMs) can perceive, ground, and utilize vertical structure information in remote-sensing natural scenes. It presents a three-stage framework—Perception, Grounding, Utilization—and a Representation Intervention Spectrum across various modalities to evaluate ten state-of-the-art models. The study finds a significant Vertical Structure Utilization Gap: while models show some geometric perception, they struggle to accurately link vertical evidence to spatial entities and incorporate it into semantic decisions.

By Jing Huang, Duanchu Wang, Junjie Yang, Zihang Cheng, Cheng Li, Lin Cui, Zhouyi Wu, Di Wang
arXiv AI
Sep 1

Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes

arXiv:2506.09557v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorat...

By Zhaoyang Wei, Bowen Jiang, Xumeng Han, Jiashu Li, Xuehui Yu, Yuling Liu, Guorong Li, Zhenjun Han, Jianbin Jiao
arXiv AI
Aug 12

Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes

arXiv:2608. 10954v1 Announce Type: cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions.

By Zhaoyang Wei, Bowen Jiang, Xumeng Han, Jiashu Li, Xuehui Yu, Yuling Liu, Guorong Li, Zhenjun Han, Jianbin Jiao
arXiv AI
Aug 26

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes

The paper introduces CAIT, a benchmark of 400 synthetic scenes featuring counter‑intuitive actions that challenge multimodal large language models (MLLMs). Human participants and proprietary models like Claude and Gemini perform well, but standard open‑source instruction‑tuned MLLMs fail, largely due to a strong language prior that overrides contradictory visual evidence. The study shows that Chain‑of‑Thought reasoning can help but introduces new issues, while targeted fine‑tuning and structured prompting can reduce reliance on language priors and improve visual grounding.

By Chen Ling, Tongwei Zhang, Hanqian Li, Nai Ding