Hugging Face Trending Papers

MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians

Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstructed environments using natural language. However, current 3D Gaussian frameworks are often bottlenecked by restrictive multiview capture requirements, costly scene-specific optimization, and the massive memory overhead of storing dense language features.

arXiv Computer Vision
Sep 24

GaussianDS: Depth-supervised Semantic Gaussian Splatting for Scene Understanding

GaussianDS introduces a depth‑supervised framework for 3D Gaussian Splatting that jointly optimizes RGB appearance, depth, and compact semantics from scratch. By arranging multi‑view images into a pose‑aware pseudo‑video and propagating view‑consistent masks via SAM2, the method aligns semantic lifting with geometric cues, using depth supervision and edge‑aware refinement to curb semantic drift and boundary leakage. The approach achieves state‑of‑the‑art performance on LERF and 3D‑OVS benchmarks while preserving high‑fidelity reconstruction and enabling downstream tasks such as 3D object removal.

By Yufei Zhang, Chenlu Zhan, Hongwei Wang
arXiv AI
Aug 25

ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian Splatting

ExtrinSplat is a new framework that separates geometry from semantics in 3D Gaussian Splatting scenes. It clusters Gaussians into overlapping 3D object groups and uses a Vision‑Language Model to generate lightweight textual hypotheses, creating an extrinsic index layer that handles complex polysemy. This approach reduces adaptation time from hours to minutes, cuts storage overhead by orders of magnitude, and outperforms existing embedding‑based methods on open‑vocabulary 3D object selection and semantic segmentation benchmarks.

By Jiayu Ding, Xinpeng Liu, Zhiyi Pan, Shiqiang Long, Ge Li
arXiv Computer Vision
Sep 21

A Scene Language Model for Open-Vocabulary Scene Mapping

The paper introduces SceneLM, a vision‑language model that maintains an open‑vocabulary 3D scene map using only a structured text list of objects as persistent memory. The model updates this textual map by adding, editing, and removing objects for each input image, learning the process through supervision tasks and an automatic annotation pipeline. Evaluations on language‑grounded retrieval and localization benchmarks show competitive performance with traditional mapping systems while producing a 6‑12× more compact representation, and the model can run online on an edge device such as a quadruped robot.

By Adam Lilja, Fabio H\"ubel, Siming He, Junsheng Fu, Claire Tomlin, Lars Hammarstrand, Jitendra Malik, Jonas Frey, Marco Pavone