PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding
arXiv:2606. 31148v1 Announce Type: cross Abstract: 3D Visual Grounding (3DVG) aims to localize target objects in 3D scenes given natural language descriptions.
The paper introduces a method that combines large language models (LLMs) with LiDAR geometry to answer complex spatial questions by grounding targets directly in LiDAR point clouds. It presents the SpatialLiDAR-QA dataset for relational grounding tasks and the SpatialLiDAR-LM model, which aligns LiDAR features with an LLM to retrieve and refine target coordinates. Experiments show significant gains over existing LiDAR–language models and multi‑camera vision‑language models in precise coordinate prediction.
arXiv:2606. 31148v1 Announce Type: cross Abstract: 3D Visual Grounding (3DVG) aims to localize target objects in 3D scenes given natural language descriptions.
arXiv:2609.06880v1 Announce Type: cross Abstract: Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial relations from a speaker's situated pers...
arXiv:2608.28762v1 Announce Type: new Abstract: Recent advances in visual question answering (VQA) and multimodal large language models (MLLMs) have enabled natural-language reasoning over traffic sc...
arXiv:2606. 17082v1 Announce Type: cross Abstract: End-to-end autonomous parking has emerged as a critical task within the realm of autonomous driving.
arXiv:2604. 08991v3 Announce Type: replace-cross Abstract: Reliable embodied interaction in indoor environments requires agents to precisely localize small everyday objects from visual observations.
arXiv:2608.13147v2 Announce Type: replace Abstract: Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera str...
arXiv:2609.17021v1 Announce Type: new Abstract: Autonomous wheel-loader control requires joint reasoning over task semantics, egocentric vision, proprioception, and 3D scene geometry. We present sens...
Autonomous wheel-loader control requires joint reasoning over task semantics, egocentric vision, proprioception, and 3D scene geometry. We present sensVLA, a Vision-Language-Action (VLA) architecture...
arXiv:2608. 16480v1 Announce Type: cross Abstract: We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vision-language reasoning in roadside sequences.
Scene coordinate regression (SCR) achieves strong performance in outdoor LiDAR localization, but it usually requires scene-specific training that can take days, limiting practical deployment. Recent w...
arXiv:2512.18991v3 Announce Type: replace-cross Abstract: Dominant paradigms for 4D LiDAR panoptic segmentation are usually required to train deep neural networks with large superimposed point clouds...
The paper introduces LiDAR-Hallu, a benchmark with 10,000 questions designed to test 4D LiDAR language models on spatio-temporal reasoning. It shows that models often achieve high multiple-choice accuracy by exploiting trivial patterns, such as always selecting the same option or relying on candidate duration, rather than truly understanding object relationships. Detailed analysis reveals systematic failures, especially in lateral-motion cases and opposite-answer scenarios, indicating that aggregate accuracy masks underlying reasoning gaps.