arXiv AI

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs

arXiv:2606. 01215v1 Announce Type: cross Abstract: Current 3D spatial reasoning methods face a fundamental trade-off: neuro-symbolic 3D (NS3D) concept learners achieve interpretable reasoning through compositional programs but are constrained to closed-set concept vocabularies and simple programs; end-to-end 3D multi-modal LLMs (3D MLLMs) could handle complex natural language and open-vocabulary concepts but suffer from black-box reasoning without explicit spatial verification.

arXiv Computation and Language
4d ago

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

arXiv:2609.38177v1 Announce Type: cross Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...

By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
arXiv Computer Vision
Aug 28

Think3D: Thinking with Space for Spatial Reasoning

Think3D introduces a framework that endows Vision‑Language Models with interactive 3D chain‑of‑thought reasoning by integrating 3D manipulation tools for active spatial exploration. The approach improves performance on benchmarks such as BLINK Multi‑view, MindCube‑1K, and VSI‑Bench‑Tiny for proprietary models like GPT‑4.1 and Gemini 2.5 Pro, and a reinforcement‑learning variant, Think3D‑RL, enables open‑weight models such as Qwen3‑VL‑4B to autonomously learn effective 3D exploration strategies, yielding tool‑use patterns comparable to stronger models and turning a performance drop on MindCube‑1K into a substantial improvement.

By Zaibin Zhang, Yuhan Wu, Lianjie Jia, Yifan Wang, Zhongbo Zhang, Yijiang Li, Binghao Ran, Fuxi Zhang, Zhuohan Sun, Yizhuang Peng, Zhenfei Yin, Lijun Wang, Huchuan Lu
arXiv AI
Jun 17

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

arXiv:2606. 17539v1 Announce Type: cross Abstract: Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging.

By Yatai Ji, An-Chieh Cheng, Yang Fu, Yukang Chen, Han Zhang, Zhaojing Yang, Wei Huang, Ka Chun Cheung, Song Han, Vidya Nariyambut Murali, Pavlo Molchanov, Jan Kautz, Simon See, Hongxu Yin, Ping Luo, Sifei Liu
Hugging Face Trending Papers
Aug 5

Disentangling 3D Modeling from Spatial Reasoning

In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perception and reasoning through large-scale training. Our key observation is that modern perception models excel at estimating continuous 3D geometry, whereas large language models (LLMs) are particularly effective at compositional and symbolic reasoning.

arXiv AI
Sep 10

MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling

MV-STRIDE is a Multi‑View hierarchical Spatial Reasoning dataset that models dependencies among perception, scene understanding, and contextual reasoning to support 3D spatial cognition. It introduces a QA generation pipeline that enforces cross‑view constraints, producing multi‑level reasoning tasks with chain‑of‑thought supervision. Experiments show that training on MV‑STRIDE yields state‑of‑the‑art performance on multi‑view spatial benchmarks, enabling MLLMs to reason robustly across diverse viewpoints.

By Jin Xu, Xiaojian Huang, Zhuodong Luo, Zhihong Zhang, Xin Liu, Jiansheng Wei, Xinzhi Wang, Jie Zhao, Xuejin Chen
arXiv Computer Vision
Sep 18

CitySTAR: Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding

CitySTAR introduces a training‑free framework that transforms billion‑scale urban point clouds into a query‑ready scene graph of open‑vocabulary 3D instances, using CodeLLM‑driven tools to supply multimodal evidence for node attributes and spatial relations. It models target‑context topology with paired hypergraphs and performs bidirectional topology verification for structural disambiguation, followed by a Reflective Cross‑modal Grounding module that integrates topology consistency and 2D visual evidence to decide over a metric‑aware 3D context graph. The authors also present CitySTAR‑3D, a benchmark that enhances semantic coverage, instance completeness, bounding‑box fidelity, and spatial‑relation complexity for city‑scale 3D grounding, and report extensive experiments showing consistent improvements in open‑world urban 3D grounding with strong interpretability and generalization.

By Shuai Zhang, Hongye Hou, Qinghe Liu, Zhuoxiao Li, Dongli Wu, Jing Ou, Yuan Liu, Wufan Zhao