arXiv AI

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

arXiv:2606. 17539v1 Announce Type: cross Abstract: Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging.

arXiv AI
Sep 10

MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling

MV-STRIDE is a Multi‑View hierarchical Spatial Reasoning dataset that models dependencies among perception, scene understanding, and contextual reasoning to support 3D spatial cognition. It introduces a QA generation pipeline that enforces cross‑view constraints, producing multi‑level reasoning tasks with chain‑of‑thought supervision. Experiments show that training on MV‑STRIDE yields state‑of‑the‑art performance on multi‑view spatial benchmarks, enabling MLLMs to reason robustly across diverse viewpoints.

By Jin Xu, Xiaojian Huang, Zhuodong Luo, Zhihong Zhang, Xin Liu, Jiansheng Wei, Xinzhi Wang, Jie Zhao, Xuejin Chen
arXiv AI
Sep 7

MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model

MultihopSpatial is a new benchmark for Vision‑Language Models that focuses on multi‑hop, compositional spatial reasoning with queries ranging from 1 to 3 hops across varied spatial perspectives. It introduces the Acc@50IoU metric, which jointly evaluates answer selection and precise bounding‑box prediction, and provides a large‑scale training corpus, MultihopSpatial‑Train, to improve spatial intelligence. Evaluation of 37 state‑of‑the‑art VLMs shows that compositional spatial reasoning remains a significant challenge, and reinforcement learning fine‑tuning on the corpus boosts both intrinsic spatial reasoning and downstream embodied manipulation performance.

By Youngwan Lee, Soojin Jang, Yoorhim Cho, Seunghwan Lee, Yong-Ju Lee, Sung Ju Hwang
arXiv Computer Vision
Sep 4

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

The paper introduces FactoSR, a factorized reinforcement learning framework designed to improve spatial reasoning in Vision‑Language Models by addressing a dimensional mismatch between 2D visual inputs and the 3D+temporal nature of the physical world. FactoSR decomposes the reasoning task into three orthogonal geometric sub‑objectives—planar correspondence (XY), depth consistency (Z), and temporal reversibility (T)—and optimizes these constraints within a unified policy learning mechanism. Experiments on multi‑view and video benchmarks show that this decomposition yields significant performance gains, achieving a 5.9% improvement on VSI‑Bench and 4.5% on All‑Angles‑Bench.

By Yijun Yang, Shenghe Zheng, Wenbo Li, Jianhui Liu, Haoze Sun, Yanbing Zhang, Jiaxiu Jiang, Lin Song, Haoyang Huang, Nan Duan, Lei Zhu
arXiv AI
Jun 12

SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

arXiv:2606. 13673v1 Announce Type: cross Abstract: Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fundamental challenge for vision-language models (VLMs).

By Seokju Cho, Ryo Hachiuma, Abhishek Badki, Hang Su, Byung-Kwan Lee, Chan Hee Song, Sifei Liu, Subhashree Radhakrishnan, Seungryong Kim, Yu-Chiang Frank Wang, Min-Hung Chen
arXiv Computer Vision
3d ago

Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation

Spatial-OPSD is a label‑free self‑improvement framework for vision‑language models that leverages spatial priors such as depth, 3D relations, and camera geometry to provide dense token‑level supervision. During training, a privileged teacher uses these priors while the student learns from only the original visual‑language input, and a recursive round‑wise scheme allows repeated self‑improvement without moving the teacher. Across four VLM families, one round of Spatial‑OPSD improves the five‑benchmark average, and three rounds push a strong spatially specialized model to the open‑source frontier, achieving the highest average among open models and best results on three of five spatial reasoning benchmarks.

By Zhenyu Liu, Zhangquan Chen, Keyi Chen, Mingze Sun, Xiang An, Haodong Jing, Ruqi Huang
arXiv AI
2d ago

Soft Spatial Reasoning

Soft Spatial Reasoning introduces a post‑training framework for Large Vision‑Language Models that replaces hard, token‑by‑token chain‑of‑thought reasoning with a soft, continuous state formed by mixing token embeddings at each intermediate step. The method employs AdaptSoft, a controller that adjusts the degree of softness based on hidden states and predictive uncertainty, guided by a gradient‑alignment learning objective that requires no intermediate supervision. Experiments on diverse spatial benchmarks show that this approach outperforms both hard and fixed‑soft chain‑of‑thought baselines and several existing LVLMs.

By Rafi Ibn Sultan, Md. Sajid Alam Chowdhury, Saleh Zare Zade, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu
arXiv AI
Sep 10

Geometry-Aware Test-Time Learning for Quantitative Spatial Reasoning

The paper introduces TTL‑SR, a geometry‑aware Test‑Time Learning framework designed to improve quantitative spatial reasoning in visual‑language models. By augmenting queries with geometrically coupled auxiliary prompts, filtering unreliable predictions, and updating models with a geometry‑aware multi‑objective loss on unlabeled test data, TTL‑SR adapts models to new domains without additional 3D supervision. Experiments show substantial accuracy gains on the Q‑Spatial‑ScanNet dataset for two state‑of‑the‑art VLMs.

By Gege Zhang, Shuaicheng Niu, Gang Dai, Lei Sun, Shuangping Huang