arXiv:2511. 07403v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial reasoning.
By Hunar Batra, Haoqin Tu, Hardy Chen, Yuanze Lin, Cihang Xie, Ronald Clark
MV-STRIDE is a Multi‑View hierarchical Spatial Reasoning dataset that models dependencies among perception, scene understanding, and contextual reasoning to support 3D spatial cognition. It introduces a QA generation pipeline that enforces cross‑view constraints, producing multi‑level reasoning tasks with chain‑of‑thought supervision. Experiments show that training on MV‑STRIDE yields state‑of‑the‑art performance on multi‑view spatial benchmarks, enabling MLLMs to reason robustly across diverse viewpoints.
By Jin Xu, Xiaojian Huang, Zhuodong Luo, Zhihong Zhang, Xin Liu, Jiansheng Wei, Xinzhi Wang, Jie Zhao, Xuejin Chen
arXiv:2606. 17539v1 Announce Type: cross Abstract: Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging.
By Yatai Ji, An-Chieh Cheng, Yang Fu, Yukang Chen, Han Zhang, Zhaojing Yang, Wei Huang, Ka Chun Cheung, Song Han, Vidya Nariyambut Murali, Pavlo Molchanov, Jan Kautz, Simon See, Hongxu Yin, Ping Luo, Sifei Liu
GraFT is a training‑free framework that enhances spatial reasoning in multimodal large language models by integrating a compact 3D scene graph (3DSG). It offers deterministic geometry via symbolic tools, allocentric layout through bird’s‑eye‑view rendering, and visual‑attribute grounding using egocentric frames. Experiments on ScanQA and VSI‑Bench show significant performance gains, with CIDEr increasing by 27% and improvements up to 65% over baseline models.
By Junqing Du, Fernando Ropero, Erkin Turkoz, Yanfeng Zhang, Lu Liu
MultihopSpatial is a new benchmark for Vision‑Language Models that focuses on multi‑hop, compositional spatial reasoning with queries ranging from 1 to 3 hops across varied spatial perspectives. It introduces the Acc@50IoU metric, which jointly evaluates answer selection and precise bounding‑box prediction, and provides a large‑scale training corpus, MultihopSpatial‑Train, to improve spatial intelligence. Evaluation of 37 state‑of‑the‑art VLMs shows that compositional spatial reasoning remains a significant challenge, and reinforcement learning fine‑tuning on the corpus boosts both intrinsic spatial reasoning and downstream embodied manipulation performance.
By Youngwan Lee, Soojin Jang, Yoorhim Cho, Seunghwan Lee, Yong-Ju Lee, Sung Ju Hwang
GaussVLA is a Vision‑Language‑Action model that enhances spatial reasoning by converting flat 2D visual tokens into compact 3D Gaussian tokens using a Gaussian Spatial Tokenizer. It further employs a Depth‑Aware Chain‑of‑Thought module to perform structured, non‑autoregressive geometric reasoning conditioned on language and flow‑time. In both simulated and real‑world tests, GaussVLA achieves high spatial‑manipulation success rates—93.5% on LIBERO and 100% on the Spatial suite—while using only 200 M parameters, outperforming SpatialVLA by 19.7% relative success.
By Md Selim Sarowar, Md Tanvir Islam, Sungho Kim, Sangtae Ahn
arXiv:2609.06880v1 Announce Type: cross
Abstract: Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial relations from a speaker's situated pers...
By Mimo Shirasaka, Haochen Zhang, Yonatan Bisk
DepthEvidence is a 4B multimodal language model that integrates dense metric depth predictions into language generation. It employs a camera‑conditioned decoder to produce full‑resolution depth maps and a dense‑to‑language interface that converts these predictions into object‑aligned geometry tokens. The model is trained with geometric supervision and instruction tuning, and it sets new state‑of‑the‑art results on a Depth‑VQA benchmark and on instance‑level metric depth estimation across nine datasets.
By Jiangning Wei, Yuan Yao, Miaomiao Cui, Mingsheng Li, Humen Zhong, Shuai Bai, Zhibo Yang
arXiv:2606. 16122v1 Announce Type: new Abstract: Visual thinking should not only sound right; it should show its evidence.
By Junkai Zhang, Yihe Deng, Kai-Wei Chang, Wei Wang
The paper introduces FactoSR, a factorized reinforcement learning framework designed to improve spatial reasoning in Vision‑Language Models by addressing a dimensional mismatch between 2D visual inputs and the 3D+temporal nature of the physical world. FactoSR decomposes the reasoning task into three orthogonal geometric sub‑objectives—planar correspondence (XY), depth consistency (Z), and temporal reversibility (T)—and optimizes these constraints within a unified policy learning mechanism. Experiments on multi‑view and video benchmarks show that this decomposition yields significant performance gains, achieving a 5.9% improvement on VSI‑Bench and 4.5% on All‑Angles‑Bench.
By Yijun Yang, Shenghe Zheng, Wenbo Li, Jianhui Liu, Haoze Sun, Yanbing Zhang, Jiaxiu Jiang, Lin Song, Haoyang Huang, Nan Duan, Lei Zhu
arXiv:2608. 05242v1 Announce Type: new Abstract: In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perception and reasoning through large-scale training.
By Haoze Sun, Jiequan Cui, Qingshan Xu, Richang Hong
arXiv:2601. 19099v2 Announce Type: replace-cross Abstract: Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views.
By Yosub Shin, Michael Buriek, Igor Molybog