arXiv AI

SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward

arXiv:2608. 12220v1 Announce Type: cross Abstract: Existing Vision-Language Models (VLMs) exhibits a critical bottleneck in robust spatial reasoning.

arXiv AI
Jun 17

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

arXiv:2606. 17539v1 Announce Type: cross Abstract: Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging.

By Yatai Ji, An-Chieh Cheng, Yang Fu, Yukang Chen, Han Zhang, Zhaojing Yang, Wei Huang, Ka Chun Cheung, Song Han, Vidya Nariyambut Murali, Pavlo Molchanov, Jan Kautz, Simon See, Hongxu Yin, Ping Luo, Sifei Liu
arXiv AI
Sep 10

MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling

MV-STRIDE is a Multi‑View hierarchical Spatial Reasoning dataset that models dependencies among perception, scene understanding, and contextual reasoning to support 3D spatial cognition. It introduces a QA generation pipeline that enforces cross‑view constraints, producing multi‑level reasoning tasks with chain‑of‑thought supervision. Experiments show that training on MV‑STRIDE yields state‑of‑the‑art performance on multi‑view spatial benchmarks, enabling MLLMs to reason robustly across diverse viewpoints.

By Jin Xu, Xiaojian Huang, Zhuodong Luo, Zhihong Zhang, Xin Liu, Jiansheng Wei, Xinzhi Wang, Jie Zhao, Xuejin Chen
arXiv AI
Jun 15

3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding

arXiv:2603. 04976v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards ( RLVR ) has emerged as a transformative paradigm for enhancing the reasoning capabilities of Large Language Models ( LLMs), yet its potential in 3D scene understanding remains under-explored.

By Xiongkun Linghu, Jiangyong Huang, Baoxiong Jia, Siyuan Huang
arXiv AI
Sep 7

MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model

MultihopSpatial is a new benchmark for Vision‑Language Models that focuses on multi‑hop, compositional spatial reasoning with queries ranging from 1 to 3 hops across varied spatial perspectives. It introduces the Acc@50IoU metric, which jointly evaluates answer selection and precise bounding‑box prediction, and provides a large‑scale training corpus, MultihopSpatial‑Train, to improve spatial intelligence. Evaluation of 37 state‑of‑the‑art VLMs shows that compositional spatial reasoning remains a significant challenge, and reinforcement learning fine‑tuning on the corpus boosts both intrinsic spatial reasoning and downstream embodied manipulation performance.

By Youngwan Lee, Soojin Jang, Yoorhim Cho, Seunghwan Lee, Yong-Ju Lee, Sung Ju Hwang
arXiv Computer Vision
Sep 4

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

The paper introduces FactoSR, a factorized reinforcement learning framework designed to improve spatial reasoning in Vision‑Language Models by addressing a dimensional mismatch between 2D visual inputs and the 3D+temporal nature of the physical world. FactoSR decomposes the reasoning task into three orthogonal geometric sub‑objectives—planar correspondence (XY), depth consistency (Z), and temporal reversibility (T)—and optimizes these constraints within a unified policy learning mechanism. Experiments on multi‑view and video benchmarks show that this decomposition yields significant performance gains, achieving a 5.9% improvement on VSI‑Bench and 4.5% on All‑Angles‑Bench.

By Yijun Yang, Shenghe Zheng, Wenbo Li, Jianhui Liu, Haoze Sun, Yanbing Zhang, Jiaxiu Jiang, Lin Song, Haoyang Huang, Nan Duan, Lei Zhu
arXiv Computer Vision
Sep 23

Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes

Metric-Bench introduces a new benchmark for Vision‑Language Models (VLMs) that focuses on metric‑spatial reasoning in indoor scenes by using in‑image reference objects with known dimensions. The accompanying MetricReasoner fine‑tuning recipe employs structured prompts and numerical rewards to implicitly learn 2D‑to‑3D mapping without camera intrinsics. Experiments show that this approach improves spatial metric understanding by 43.1 % over existing models and boosts downstream embodied tasks, while also delivering gains on general VLM benchmarks.

By Yuling Xi, Haokai Zhang, Muzhi Zhu, Hao Zhong, Zongze Du, Hengyu Zhao, Chenchen Jing, Yufei Yin, Bin Qin, Yongjie Yang, Zhenbo Luo, Hao Chen, Chunhua Shen
Hugging Face Trending Papers
Sep 3

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

The paper introduces FactoSR, a factorized reinforcement learning framework designed to improve spatial reasoning in Vision‑Language Models (VLMs). By decomposing the problem into planar correspondence (XY), depth consistency (Z), and temporal reversibility (T), FactoSR addresses the dimensional mismatch between 2D visual inputs and 3D physical reasoning. Experiments on multi‑view and video benchmarks show significant performance gains, with a 5.9% improvement on VSI‑Bench and 4.5% on All‑Angles‑Bench.