arXiv:2603. 06576v2 Announce Type: replace-cross Abstract: The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail scenarios.
By Thomas Monninger, Shaoyuan Xie, Qi Alfred Chen, Sihao Ding
arXiv:2608.21136v1 Announce Type: new
Abstract: Recently, open-vocabulary zero-shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data-intensive supe...
By Jie Xu, Na Zhao
Lang3DSeg introduces a point‑transformer backbone for open‑vocabulary, annotation‑free 3D LiDAR segmentation, trained from scratch without geometric pre‑training. It tackles noise from 2D‑to‑3D label projections by applying a class‑priority rule and truncating projected instances at depth gaps, thereby correcting depth‑ambiguity errors. The method achieves state‑of‑the‑art results on nuScenes (52.8 % mIoU) and SemanticKITTI (41.4 % mIoU) while operating in real‑time on a single LiDAR sweep.
By Cigdem Kokenoz, Amir Salarpour, Alkim Domeke, Christopher Salas, Pedram MohajerAnsari, Long Cheng, Mert D. Pes\'e, Bing Li
PointGauss is a 3D-native framework that performs semantic parsing and instance segmentation on 3D Gaussian splatting representations by treating Gaussian primitives as unstructured point sets and extracting scale‑invariant geometric features with Point Transformer V3. It introduces an adaptive region‑of‑interest cropping strategy and an instance‑aware distance‑constrained rasterization pipeline to enable scalable, view‑consistent pixel‑level projections. The authors also release SplatSeg‑360, a cross‑scale benchmark with 32 complex scenes and over 6,300 aligned 2D‑3D masks, and show that PointGauss achieves real‑time performance with state‑of‑the‑art 3D‑mIoU (~90%) and 2D‑mIoU (~80%) scores.
By Wentao Sun, Yiping Chen, John S. Zelek, Jonathan Li
arXiv:2608.13147v2 Announce Type: replace
Abstract: Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera str...
By Longfei Xu, Xiaohui Wang, Zehao Huang, Han Li, Ya Yang, Naiyan Wang, Si Liu
SAM‑V is a geometry‑aware extension of the Segment Anything Model (SAM) that integrates 3D priors from a feed‑forward geometry model (VGGT) into 2D segmentation. It uses a prompt‑fusion mechanism to combine sparse SAM prompts with view‑specific camera tokens and local VGGT features, enabling a mask decoder that attends to both dense 2D and 3D cues. The resulting end‑to‑end system produces consistent multi‑view instance segmentation in a single forward pass, achieving significant gains on the IGGT 3D tracking benchmark without offline mask matching or explicit 3D reconstruction.
By Jiangshan Gong, Yuqun Wu, Qiqian Fu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem