MV-STRIDE is a Multi‑View hierarchical Spatial Reasoning dataset that models dependencies among perception, scene understanding, and contextual reasoning to support 3D spatial cognition. It introduces a QA generation pipeline that enforces cross‑view constraints, producing multi‑level reasoning tasks with chain‑of‑thought supervision. Experiments show that training on MV‑STRIDE yields state‑of‑the‑art performance on multi‑view spatial benchmarks, enabling MLLMs to reason robustly across diverse viewpoints.
By Jin Xu, Xiaojian Huang, Zhuodong Luo, Zhihong Zhang, Xin Liu, Jiansheng Wei, Xinzhi Wang, Jie Zhao, Xuejin Chen
arXiv:2609.36219v1 Announce Type: new
Abstract: Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as...
By Bang Xiao, Wenqi Jia, Ozgur Kara, Tiancheng Shen, Yibo Yang, Bolin Lai, Junho Kim, James Matthew Rehg
GraFT is a training‑free framework that enhances spatial reasoning in multimodal large language models by integrating a compact 3D scene graph (3DSG). It offers deterministic geometry via symbolic tools, allocentric layout through bird’s‑eye‑view rendering, and visual‑attribute grounding using egocentric frames. Experiments on ScanQA and VSI‑Bench show significant performance gains, with CIDEr increasing by 27% and improvements up to 65% over baseline models.
By Junqing Du, Fernando Ropero, Erkin Turkoz, Yanfeng Zhang, Lu Liu
arXiv:2609.38177v1 Announce Type: cross
Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...
By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
arXiv:2605. 16713v2 Announce Type: replace-cross Abstract: Modern Vision-Language Models (VLMs) achieve strong semantic recognition, yet remain brittle on elementary spatial relations such as left of, on, behind, and between.
By Renjie Gu, Kaichen Zhou, Yan Luo, Mengyu Wang
arXiv:2609.24813v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) have demonstrated remarkable capabilities in multimodal tasks, yet they still exhibit poor ability in spatial reasoning....
By Haoran Sun, Jingqi Xu, Yanhui Li, Enci Liu, Kaidi Xu, Yanwei Liu