Hugging Face Trending Papers

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

Read the original on Hugging Face Trending Papers →

As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitoring, the opacity of their spatial judgments limits operator trust and auditability. MLLMs demonstrate strong reasoning but often struggle with fine-grained spatial understanding and object hallucination.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Jun 17

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

arXiv:2606. 17539v1 Announce Type: cross Abstract: Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging.

By Yatai Ji, An-Chieh Cheng, Yang Fu, Yukang Chen, Han Zhang, Zhaojing Yang, Wei Huang, Ka Chun Cheung, Song Han, Vidya Nariyambut Murali, Pavlo Molchanov, Jan Kautz, Simon See, Hongxu Yin, Ping Luo, Sifei Liu
arXiv AI
6d ago

DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models

DepthEvidence is a 4B multimodal language model that integrates dense metric depth predictions into language generation. It employs a camera‑conditioned decoder to produce full‑resolution depth maps and a dense‑to‑language interface that converts these predictions into object‑aligned geometry tokens. The model is trained with geometric supervision and instruction tuning, and it sets new state‑of‑the‑art results on a Depth‑VQA benchmark and on instance‑level metric depth estimation across nine datasets.

By Jiangning Wei, Yuan Yao, Miaomiao Cui, Mingsheng Li, Humen Zhong, Shuai Bai, Zhibo Yang
arXiv Computer Vision
2d ago

Beyond Localization: A Comprehensive Benchmark of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

The paper introduces PCSR-Bench, a benchmark of 84,373 question‑answer pairs derived from 2,600 omnidirectional images across 26 indoor environments, designed to evaluate perspective‑conditioned spatial reasoning (PCSR) in multimodal large language models (MLLMs). It reports a significant perception–reasoning gap, with accuracy dropping from 57.59% on limited field‑of‑view reasoning to as low as 0.64% on open‑ended compositional directional chains. An RL‑based diagnostic study on a 7B‑scale model shows that reward shaping can improve performance to 60.06% on a controlled task, indicating partial plasticity of PCSR capabilities.

By Yuangong Chen, Wai Keung Wong, Jiaxing Li, Ioannis Patras, Xu Zheng