arXiv AI

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning

arXiv:2602. 08236v2 Announce Type: replace-cross Abstract: Despite rapid progress in MLLMs, visual spatial reasoning remains unreliable when correct answers depend on how a scene would appear under unseen or alternative viewpoints.

arXiv AI
Jun 9

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

arXiv:2606. 09669v1 Announce Type: new Abstract: Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world.

By Hongcheng Gao, Hailong Qu, Jingyi Tang, Jiahao Wang, Zihao Huang, Hengkang Qiao, Shihong Huang, Junming Yang, Yi Li, Hongyixuan Yuan, Wenjie Li, Bohan Zeng, Wenbo Li, Bo Wang, Jianhui Liu, Olive Huang, Haoyang Huang, Wentao Zhang, Guoqing Huang, Nan Duan, Yinpeng Dong
arXiv Machine Learning
Aug 21

Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

arXiv:2608. 19669v1 Announce Type: cross Abstract: Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage.

By Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi
arXiv AI
Aug 24

Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds

The paper investigates how visual presentation affects vision‑language models (VLMs) on the SPaRC spatial planning benchmark. By adding lightweight input‑side scaffolds that keep the visual modality but make spatial structure clearer, the authors achieve up to a 34.0‑percentage‑point accuracy boost across multiple VLMs, and an additional 4.6 points when combined with GRPO training. Analyses reveal that these improvements stem mainly from reduced grounding errors, while rule‑based reasoning remains difficult, highlighting visual presentation as a key determinant of whether VLM benchmarks test grounded perception, downstream reasoning, or both.

By Lars Benedikt Kaesberg, Tianyu Yang, Florian Valentin Wunderlich, Terry Ruas, Jan Philip Wahle, Daniel Kurzawe, Bela Gipp