Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clo...
arXiv:2609.16233v1 Announce Type: cross
Abstract: Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current ben...
By Anubhav Khanal, Prabigya Acharya, Roshni Poudel, Sujan Kapali, Bigyan Bhatta, Pramish Paudel, Francois Rameau, Danda Pani Paudel
Previous work has evaluated physics reasoning in foundation models using synthetic or semi-synthetic scenes and visual question-answering tasks. However, these benchmarks emphasize high-level events and lack the visual fidelity required to assess true low-level Newtonian understanding.
The paper introduces SciGram, a large-scale dataset of 194K scientific diagrams paired with 1.4M visual instructions generated through a terminology‑grounded pipeline that extracts domain concepts, synthesizes facts, and retrieves relevant diagrams. Models fine‑tuned on SciGram show significant gains on diagram‑centric benchmarks such as TQA, ScienceQA, and AI2D, and when combined with existing models like LLaVA OneVision, set new state‑of‑the‑art performance. The authors release both the dataset and trained models to support further research in scientific diagram understanding.
By Raul Ortega, Jos\'e Manuel G\'omez-P\'erez
While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous semantic alignment and logical reasoning required for scientific imagery. Inspired by Peirce's Semiotic Triad, we introduce Scientific Image Reasoning (SciIR), a comprehensive resource for training and evaluation of scientific image generation.
arXiv:2606. 29786v2 Announce Type: replace Abstract: 3D scene graphs (3DSGs) provide a compact and structured abstraction of 3D environments.
By Yirum Kim, Ue-Hwan Kim
arXiv:2602. 08058v3 Announce Type: replace-cross Abstract: In the presence of occlusions and measurement noise, geometrically accurate scene reconstructions -- which fit the sensor data -- can still be physically incorrect.
By Xihang Yu, Rajat Talak, Lorenzo Shaikewitz, Luca Carlone
GraFT is a training‑free framework that enhances spatial reasoning in multimodal large language models by integrating a compact 3D scene graph (3DSG). It offers deterministic geometry via symbolic tools, allocentric layout through bird’s‑eye‑view rendering, and visual‑attribute grounding using egocentric frames. Experiments on ScanQA and VSI‑Bench show significant performance gains, with CIDEr increasing by 27% and improvements up to 65% over baseline models.
By Junqing Du, Fernando Ropero, Erkin Turkoz, Yanfeng Zhang, Lu Liu
arXiv:2601. 11729v2 Announce Type: replace-cross Abstract: Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicability to embodied systems.
By Turhan Can Kargin, Wojciech Jasi\'nski, Adam Pardyl, Bartosz Zieli\'nski, Marcin Przewi\k{e}\'zlikowski
arXiv:2606. 29667v1 Announce Type: cross Abstract: The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains locked away and inaccessible to AI at scale.
By Subham Ghosh, Shubham Tiwari, Mohammad Ibrahim, Abhishek Tewari
SnapPhysics is a training‑free framework that reconstructs 3D objects and estimates their physical properties—such as mass, friction, and center of gravity—from a single image. It combines instance‑level 3D reconstruction with a physics‑aware scene graph that encodes inter‑object relationships, providing structured context for vision‑language model reasoning. Experiments on 3D‑FRONT and real captured scenes show significant improvements over prior methods, reducing errors in mass estimation and enhancing scene‑level F‑Score.
arXiv:2606. 07529v1 Announce Type: cross Abstract: Large language models (LLMs) have recently been applied to 3D vision-language (3D-VL) tasks, which require spatial reasoning to identify target objects relative to anchors.
By Shengli Zhou, Xiangchen Wang, Guanhua Chen, Feng Zheng