arXiv Computer Vision By Runyi Yang, Murat Akkoyun, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Xiaoye Wang, Kailun Yang, Danda Pani Paudel, Luc Van Gool, Kunyu Peng

Do LiDAR Language Models Really Understand Spatio-temporal Relationships?

Read the original on arXiv Computer Vision →

The paper introduces LiDAR-Hallu, a benchmark with 10,000 questions designed to test 4D LiDAR language models on spatio-temporal reasoning. It shows that models often achieve high multiple-choice accuracy by exploiting trivial patterns, such as always selecting the same option or relying on candidate duration, rather than truly understanding object relationships. Detailed analysis reveals systematic failures, especially in lateral-motion cases and opposite-answer scenarios, indicating that aggregate accuracy masks underlying reasoning gaps.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computation and Language
Sep 25

STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models

STRAND is a new benchmark that tests multimodal large language models’ ability to track objects, their states, and relationships over time in videos. It evaluates intermediate reasoning by breaking queries into sub‑questions and uses Faithful Accuracy to ensure all parts of an answer are correct. The authors also propose an object‑centric framework that builds structured trajectories and shows reduced hallucinations and better temporal consistency compared to existing models.

By Thong Nguyen, Tri Cao, Khoi Le, Cong-Duy Nguyen, Quynh Vo, See-Kiong Ng, Bryan Hooi Kuen-Yew
arXiv Computer Vision
Sep 25

Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding

The paper introduces a method that combines large language models (LLMs) with LiDAR geometry to answer complex spatial questions by grounding targets directly in LiDAR point clouds. It presents the SpatialLiDAR-QA dataset for relational grounding tasks and the SpatialLiDAR-LM model, which aligns LiDAR features with an LLM to retrieve and refine target coordinates. Experiments show significant gains over existing LiDAR–language models and multi‑camera vision‑language models in precise coordinate prediction.

By Byounggun Park, Giyong Moon, Jusung Kim, Soonmin Hwang