STRAND is a new benchmark that tests multimodal large language models’ ability to track objects, their states, and relationships over time in videos. It evaluates intermediate reasoning by breaking queries into sub‑questions and uses Faithful Accuracy to ensure all parts of an answer are correct. The authors also propose an object‑centric framework that builds structured trajectories and shows reduced hallucinations and better temporal consistency compared to existing models.
By Thong Nguyen, Tri Cao, Khoi Le, Cong-Duy Nguyen, Quynh Vo, See-Kiong Ng, Bryan Hooi Kuen-Yew
The paper introduces a method that combines large language models (LLMs) with LiDAR geometry to answer complex spatial questions by grounding targets directly in LiDAR point clouds. It presents the SpatialLiDAR-QA dataset for relational grounding tasks and the SpatialLiDAR-LM model, which aligns LiDAR features with an LLM to retrieve and refine target coordinates. Experiments show significant gains over existing LiDAR–language models and multi‑camera vision‑language models in precise coordinate prediction.
By Byounggun Park, Giyong Moon, Jusung Kim, Soonmin Hwang
arXiv:2607. 00491v1 Announce Type: cross Abstract: Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input.
By Leyuan Yu, Xiao Tang, Minghao Liu, Xinyuan Li, Xiaokai Bai, Sheng Zhou, Qunshu Lin, Weihao Xuan, Naoto Yokoya
arXiv:2605. 08974v2 Announce Type: replace-cross Abstract: While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes.
By Tri Cao, Khoi Le, Thong Nguyen, Cong-Duy Nguyen, Quynh Vo, Anh Tuan Luu, Chunyan Miao, See-Kiong Ng, Shuicheng Yan, Bryan Hooi
arXiv:2609.39969v1 Announce Type: cross
Abstract: Physical LiDAR attacks are often evaluated using fixed primitives and manually selected parameters, despite their strong dependence on surrounding tr...
By Yiming Gao, Shaocheng Luo
arXiv:2608. 19522v1 Announce Type: cross Abstract: Scan-to-map LiDAR odometry drifts unboundedly along the unobservable axes of geometrically degenerate environments like tunnels and corridors, and existing degeneracy handling requires environment-specific parameter tuning.
By Eunsoo Im