arXiv Machine Learning By Animesh Maheshwari, Divyansh Sahu, Nishit Verma

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects?

Read the original on arXiv Machine Learning →

arXiv:2605. 20448v2 Announce Type: replace-cross Abstract: Vision-language models reliably name objects in a scene, but do they represent the 3D layout those objects inhabit?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
4d ago

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

arXiv:2609.38177v1 Announce Type: cross Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...

By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
Hugging Face Trending Papers
Jul 29

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitoring, the opacity of their spatial judgments limits operator trust and auditability. MLLMs demonstrate strong reasoning but often struggle with fine-grained spatial understanding and object hallucination.