GraFT is a training‑free framework that enhances spatial reasoning in multimodal large language models by integrating a compact 3D scene graph (3DSG). It offers deterministic geometry via symbolic tools, allocentric layout through bird’s‑eye‑view rendering, and visual‑attribute grounding using egocentric frames. Experiments on ScanQA and VSI‑Bench show significant performance gains, with CIDEr increasing by 27% and improvements up to 65% over baseline models.
By Junqing Du, Fernando Ropero, Erkin Turkoz, Yanfeng Zhang, Lu Liu
arXiv:2609.38177v1 Announce Type: cross
Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...
By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
MV-STRIDE is a Multi‑View hierarchical Spatial Reasoning dataset that models dependencies among perception, scene understanding, and contextual reasoning to support 3D spatial cognition. It introduces a QA generation pipeline that enforces cross‑view constraints, producing multi‑level reasoning tasks with chain‑of‑thought supervision. Experiments show that training on MV‑STRIDE yields state‑of‑the‑art performance on multi‑view spatial benchmarks, enabling MLLMs to reason robustly across diverse viewpoints.
By Jin Xu, Xiaojian Huang, Zhuodong Luo, Zhihong Zhang, Xin Liu, Jiansheng Wei, Xinzhi Wang, Jie Zhao, Xuejin Chen
arXiv:2511. 07403v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial reasoning.
By Hunar Batra, Haoqin Tu, Hardy Chen, Yuanze Lin, Cihang Xie, Ronald Clark
arXiv:2610.02021v1 Announce Type: cross
Abstract: Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by da...
By Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel
arXiv:2609.15137v1 Announce Type: cross
Abstract: 3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answering (VQA), but their dense semantic f...
By Davit Soselia, Joseph JaJa, Amitabh Varshney