arXiv Computer Vision

Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers

Lang3DSeg introduces a point‑transformer backbone for open‑vocabulary, annotation‑free 3D LiDAR segmentation, trained from scratch without geometric pre‑training. It tackles noise from 2D‑to‑3D label projections by applying a class‑priority rule and truncating projected instances at depth gaps, thereby correcting depth‑ambiguity errors. The method achieves state‑of‑the‑art results on nuScenes (52.8 % mIoU) and SemanticKITTI (41.4 % mIoU) while operating in real‑time on a single LiDAR sweep.

arXiv Computer Vision
Aug 27

Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models

The paper introduces LiDAR‑SAM2, a framework that converts the 2D video foundation model SAM2 into a scalable source of supervision for 4D LiDAR data. By projecting SAM2 video masks into multi‑view LiDAR space and aggregating them temporally, the method automatically generates temporally coherent LiDAR labels without human annotation. Experiments on SemanticKITTI show that these automatically produced semantic and panoptic labels achieve quality close to full human annotation, enabling models trained on them to approach the performance of fully supervised systems.

By Jihun Kim, Hyun-Kurl Jang, Hyemin Yang, Jinnyeong Yang, Hyeokjun Kweon, Kuk-Jin Yoon
arXiv Computer Vision
Sep 25

SplatLabel: Pseudo-Labelling through 4D Gaussian Splatting

SplatLabel is an automated pipeline that uses a 4D Gaussian representation to generate LiDAR segmentation and semantic occupancy grids with predictive confidence. It models dynamic scenes through an explicit temporal manifold, tracking moving actors without requiring pre‑annotated 3D bounding boxes. By integrating 360‑degree LiDAR depth maps and distilling soft probabilities from 2D models, it resolves semantic ambiguities over time and space, and evaluates pseudo‑labels via a selective classification framework that balances precision and recall.

By Nitya Nanvani, Andras Palffy, Holger Caesar
arXiv Computer Vision
Sep 18

Open-vocabulary 3D object detection with promptable segmentation

The paper introduces an open‑vocabulary 3D object detection pipeline that uses a promptable segmentation model (SAM3) to generate instance masks from six surround‑view cameras. These masks are converted into metric 3D boxes, achieving up to 0.413 mAP/0.555 NDS without any training when supervised box geometry is borrowed at inference. The approach also improves a supervised LiDAR‑only detector by 0.034 mAP through a camera‑witness rule, demonstrating that measurement precision, not 2D detection, limits performance.

By \"Omer Faruk Deniz, Mustafa Taha Ko\c{c}yi\u{g}it
arXiv AI
Sep 10

GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

GoDeep is an annotation‑free method for open‑vocabulary 3D scene understanding that uses a vision‑language model solely as a translator to generate structured, entity‑level descriptions of each image. These descriptions are projected and aggregated in a language‑only embedding space, eliminating the need for a 3D training corpus or domain‑specific encoder. The approach achieves competitive performance on ScanNet++ and a cultural heritage benchmark, accurately localizes out‑of‑vocabulary objects, and offers explainable, point‑level predictions.

By Thodoris Betsas, Anastasios Doulamis, Andreas Georgopoulos
arXiv Computer Vision
Sep 18

SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation

SenseFuse introduces a label‑free fusion approach that balances 2D image and 3D shape encoders for open‑vocabulary 3D instance segmentation. By selecting a scene‑level fusion weight through an adaptive, sensitivity‑based mechanism, it improves mask labeling accuracy across multiple datasets, recovering up to 93% of the potential gain from an oracle weight. The method demonstrates that image and shape encoders have complementary failure patterns, leading to higher instance AP in most evaluated settings.

By Euiseok Han, Tri Ton, Hwanhee Kim, Seungyeon Ryu, Chang D. Yoo