arXiv Computer Vision

Beyond Localization: A Comprehensive Benchmark of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

The paper introduces PCSR-Bench, a benchmark of 84,373 question‑answer pairs derived from 2,600 omnidirectional images across 26 indoor environments, designed to evaluate perspective‑conditioned spatial reasoning (PCSR) in multimodal large language models (MLLMs). It reports a significant perception–reasoning gap, with accuracy dropping from 57.59% on limited field‑of‑view reasoning to as low as 0.64% on open‑ended compositional directional chains. An RL‑based diagnostic study on a 7B‑scale model shows that reward shaping can improve performance to 60.06% on a controlled task, indicating partial plasticity of PCSR capabilities.

arXiv AI
Jun 17

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

arXiv:2606. 17539v1 Announce Type: cross Abstract: Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging.

By Yatai Ji, An-Chieh Cheng, Yang Fu, Yukang Chen, Han Zhang, Zhaojing Yang, Wei Huang, Ka Chun Cheung, Song Han, Vidya Nariyambut Murali, Pavlo Molchanov, Jan Kautz, Simon See, Hongxu Yin, Ping Luo, Sifei Liu
arXiv AI
Aug 24

StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models

StateSight is a new benchmark designed to isolate and evaluate the ability of vision‑language models to reconstruct latent spatial structure from a single image. It consists of three procedurally generated task families—cube‑net opposite‑face reasoning, occluded cube‑tower counting, and 4‑neighbor connected‑component counting—each with 300 deterministic prompts and exact‑match scoring. The benchmark also includes a companion dataset, StateSight‑Steps, with 900 image‑text examples and 3,600 intermediate visual states to aid analysis of reconstruction errors.

By Michelle Lin
arXiv Computer Vision
Aug 27

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs

OVO‑S‑Bench is a fully human‑annotated benchmark designed to evaluate streaming spatial intelligence in multimodal large language models (MLLMs). It contains 1,680 questions derived from 348 source videos, each with a query timestamp and evidence interval, and tests models on four levels of abstraction: instantaneous egocentric perception, spatiotemporal context tracking, generative spatial reasoning, and allocentric spatial mapping. Across 38 MLLMs, Gemini‑3.1‑Pro scored 59.2 versus 92.2 for human experts, with allocentric spatial mapping identified as the main challenge, and the benchmark reveals that chain‑of‑thought reasoning can worsen spatial errors when not grounded in the stream.

By Yifei Li, Pengyiang Liu, Yuhang Zang, Zhongyue Shi, Qi Fu, Hongye Hao, Jiwen Lu
arXiv AI
Jun 9

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

arXiv:2606. 09669v1 Announce Type: new Abstract: Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world.

By Hongcheng Gao, Hailong Qu, Jingyi Tang, Jiahao Wang, Zihao Huang, Hengkang Qiao, Shihong Huang, Junming Yang, Yi Li, Hongyixuan Yuan, Wenjie Li, Bohan Zeng, Wenbo Li, Bo Wang, Jianhui Liu, Olive Huang, Haoyang Huang, Wentao Zhang, Guoqing Huang, Nan Duan, Yinpeng Dong