arXiv AI By Yang Zhou, Zixuan Huang, Sunzhu Li, Zhuo Yang, Chen Zhang, Shunian Chen, Caijun Yan, Jianyao Xu, Shunyu Liu, Weijie Fu, Peiliang Li, Xiaozhi Chen, Yuxiang Cai

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

Read the original on arXiv AI →

arXiv:2607. 27703v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 28

Think3D: Thinking with Space for Spatial Reasoning

Think3D introduces a framework that endows Vision‑Language Models with interactive 3D chain‑of‑thought reasoning by integrating 3D manipulation tools for active spatial exploration. The approach improves performance on benchmarks such as BLINK Multi‑view, MindCube‑1K, and VSI‑Bench‑Tiny for proprietary models like GPT‑4.1 and Gemini 2.5 Pro, and a reinforcement‑learning variant, Think3D‑RL, enables open‑weight models such as Qwen3‑VL‑4B to autonomously learn effective 3D exploration strategies, yielding tool‑use patterns comparable to stronger models and turning a performance drop on MindCube‑1K into a substantial improvement.

By Zaibin Zhang, Yuhan Wu, Lianjie Jia, Yifan Wang, Zhongbo Zhang, Yijiang Li, Binghao Ran, Fuxi Zhang, Zhuohan Sun, Yizhuang Peng, Zhenfei Yin, Lijun Wang, Huchuan Lu
arXiv AI
Jun 12

SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

arXiv:2606. 13673v1 Announce Type: cross Abstract: Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fundamental challenge for vision-language models (VLMs).

By Seokju Cho, Ryo Hachiuma, Abhishek Badki, Hang Su, Byung-Kwan Lee, Chan Hee Song, Sifei Liu, Subhashree Radhakrishnan, Seungryong Kim, Yu-Chiang Frank Wang, Min-Hung Chen
arXiv Computer Vision
2d ago

Beyond Localization: A Comprehensive Benchmark of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

The paper introduces PCSR-Bench, a benchmark of 84,373 question‑answer pairs derived from 2,600 omnidirectional images across 26 indoor environments, designed to evaluate perspective‑conditioned spatial reasoning (PCSR) in multimodal large language models (MLLMs). It reports a significant perception–reasoning gap, with accuracy dropping from 57.59% on limited field‑of‑view reasoning to as low as 0.64% on open‑ended compositional directional chains. An RL‑based diagnostic study on a 7B‑scale model shows that reward shaping can improve performance to 60.06% on a controlled task, indicating partial plasticity of PCSR capabilities.

By Yuangong Chen, Wai Keung Wong, Jiaxing Li, Ioannis Patras, Xu Zheng