arXiv Computer Vision

TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking

TRACKGRAPH is an online open‑vocabulary 3D mapping system that tracks 2D masks in the image stream before fusing them into a class‑agnostic 3D segment layer within a hierarchical scene graph. It uses FastSAM and CLIP for sparse keyframes, DINOv3 for dense mask propagation, and compact multi‑view CLIP embeddings for open‑vocabulary retrieval. The method outperforms state‑of‑the‑art mapping techniques on Replica, ScanNet++, and HM3D, achieving higher synonym frequency, faster processing, and lower GPU memory usage, and has been deployed on quadruped robots and drones at real‑time rates.

arXiv Computer Vision
Sep 21

A Scene Language Model for Open-Vocabulary Scene Mapping

The paper introduces SceneLM, a vision‑language model that maintains an open‑vocabulary 3D scene map using only a structured text list of objects as persistent memory. The model updates this textual map by adding, editing, and removing objects for each input image, learning the process through supervision tasks and an automatic annotation pipeline. Evaluations on language‑grounded retrieval and localization benchmarks show competitive performance with traditional mapping systems while producing a 6‑12× more compact representation, and the model can run online on an edge device such as a quadruped robot.

By Adam Lilja, Fabio H\"ubel, Siming He, Junsheng Fu, Claire Tomlin, Lars Hammarstrand, Jitendra Malik, Jonas Frey, Marco Pavone
arXiv AI
Sep 10

GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

GoDeep is an annotation‑free method for open‑vocabulary 3D scene understanding that uses a vision‑language model solely as a translator to generate structured, entity‑level descriptions of each image. These descriptions are projected and aggregated in a language‑only embedding space, eliminating the need for a 3D training corpus or domain‑specific encoder. The approach achieves competitive performance on ScanNet++ and a cultural heritage benchmark, accurately localizes out‑of‑vocabulary objects, and offers explainable, point‑level predictions.

By Thodoris Betsas, Anastasios Doulamis, Andreas Georgopoulos
Hugging Face Trending Papers
Jul 30

MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians

Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstructed environments using natural language. However, current 3D Gaussian frameworks are often bottlenecked by restrictive multiview capture requirements, costly scene-specific optimization, and the massive memory overhead of storing dense language features.

arXiv Computer Vision
Sep 18

SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation

SenseFuse introduces a label‑free fusion approach that balances 2D image and 3D shape encoders for open‑vocabulary 3D instance segmentation. By selecting a scene‑level fusion weight through an adaptive, sensitivity‑based mechanism, it improves mask labeling accuracy across multiple datasets, recovering up to 93% of the potential gain from an oracle weight. The method demonstrates that image and shape encoders have complementary failure patterns, leading to higher instance AP in most evaluated settings.

By Euiseok Han, Tri Ton, Hwanhee Kim, Seungyeon Ryu, Chang D. Yoo
arXiv AI
Sep 25

TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations

TrackEverything is a 3D point tracker that overcomes the trade‑off between sparse long‑horizon tracking and dense short‑clip tracking by representing videos as persistent 3D scene tracks in world coordinates. It introduces voxel‑based de‑duplication at sliding‑window boundaries, a two‑stage refinement process (endpoint refiner and lightweight trajectory refiner), and a 3D WAFT module that replaces memory‑heavy 4D correlation volumes with efficient feature sampling. The method can track all visible points in videos longer than 1000 frames using only 40 GB of GPU memory, outperforming existing dense trackers on short clips and matching sparse trackers on long sequences.

By Ayush Jain, Sreeharsha Paruchuri, Ishita Gupta, Fan Zhang, Tanner Schmidt, Jakob Engel, Katerina Fragkiadaki, Adam W. Harley