arXiv:2607. 13454v1 Announce Type: cross Abstract: Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge.
By Hao Li, Han Fang, Zixin Pan, Xin Wei, Hongbo Sun, Jinglin Xu, Zhiyu Lin, Ye Yuan, Zhongjiang He, Yu Yu, Hao Sun
Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge. Existing methods primarily rely on symbolic text tokens, which inherently lack the fidelity to represent continuous geometric information.
Spatial-OPSD is a label‑free self‑improvement framework for vision‑language models that leverages spatial priors such as depth, 3D relations, and camera geometry to provide dense token‑level supervision. During training, a privileged teacher uses these priors while the student learns from only the original visual‑language input, and a recursive round‑wise scheme allows repeated self‑improvement without moving the teacher. Across four VLM families, one round of Spatial‑OPSD improves the five‑benchmark average, and three rounds push a strong spatially specialized model to the open‑source frontier, achieving the highest average among open models and best results on three of five spatial reasoning benchmarks.
By Zhenyu Liu, Zhangquan Chen, Keyi Chen, Mingze Sun, Xiang An, Haodong Jing, Ruqi Huang
The paper introduces TTL‑SR, a geometry‑aware Test‑Time Learning framework designed to improve quantitative spatial reasoning in visual‑language models. By augmenting queries with geometrically coupled auxiliary prompts, filtering unreliable predictions, and updating models with a geometry‑aware multi‑objective loss on unlabeled test data, TTL‑SR adapts models to new domains without additional 3D supervision. Experiments show substantial accuracy gains on the Q‑Spatial‑ScanNet dataset for two state‑of‑the‑art VLMs.
By Gege Zhang, Shuaicheng Niu, Gang Dai, Lei Sun, Shuangping Huang
A*-Thought-V2 is a framework that models Chain-of-Thought reasoning as a geometric trajectory in a 3D PCA space, using explicit-implicit latent tokens to compress steps that deviate from the main question-to-solution direction. The method measures alignment angles to decide which steps remain text and which become latent, and introduces stepwise embedding forcing and label forcing to train the architecture. Experiments on Qwen models show up to 2.6% accuracy gains, halved response length, and significant reductions in computation and training time.
By Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He
arXiv:2606. 26535v1 Announce Type: cross Abstract: Current VLM evaluations often conflate language priors with genuine spatial reasoning.
By Zhixing Li, Yinan Yu
Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit perception and explicit reasoning.
arXiv:2511. 07403v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial reasoning.
By Hunar Batra, Haoqin Tu, Hardy Chen, Yuanze Lin, Cihang Xie, Ronald Clark
The paper introduces FactoSR, a factorized reinforcement learning framework designed to improve spatial reasoning in Vision‑Language Models (VLMs). By decomposing the problem into planar correspondence (XY), depth consistency (Z), and temporal reversibility (T), FactoSR addresses the dimensional mismatch between 2D visual inputs and 3D physical reasoning. Experiments on multi‑view and video benchmarks show significant performance gains, with a 5.9% improvement on VSI‑Bench and 4.5% on All‑Angles‑Bench.
arXiv:2605.18641v2 Announce Type: replace
Abstract: Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generat...
By Dongyao Zhu, Zhen Wang, Xi Xiao, Han Jiang, Saeed Vahidian, Wei-Lun Chao, Tanya Berger-Wolf, Yu Su, Raju Vatsavai, Jianyang Gu
arXiv:2608. 19669v1 Announce Type: cross Abstract: Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage.
By Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi
Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage. In this paper, we identify two key limitations of this framework, one in each stage.