The paper introduces SceneLM, a vision‑language model that maintains an open‑vocabulary 3D scene map using only a structured text list of objects as persistent memory. The model updates this textual map by adding, editing, and removing objects for each input image, learning the process through supervision tasks and an automatic annotation pipeline. Evaluations on language‑grounded retrieval and localization benchmarks show competitive performance with traditional mapping systems while producing a 6‑12× more compact representation, and the model can run online on an edge device such as a quadruped robot.
By Adam Lilja, Fabio H\"ubel, Siming He, Junsheng Fu, Claire Tomlin, Lars Hammarstrand, Jitendra Malik, Jonas Frey, Marco Pavone
GoDeep is an annotation‑free method for open‑vocabulary 3D scene understanding that uses a vision‑language model solely as a translator to generate structured, entity‑level descriptions of each image. These descriptions are projected and aggregated in a language‑only embedding space, eliminating the need for a 3D training corpus or domain‑specific encoder. The approach achieves competitive performance on ScanNet++ and a cultural heritage benchmark, accurately localizes out‑of‑vocabulary objects, and offers explainable, point‑level predictions.
By Thodoris Betsas, Anastasios Doulamis, Andreas Georgopoulos
arXiv:2509. 24528v4 Announce Type: replace-cross Abstract: Object retrieval from a scene has become a new trend of research due to its numerous applications.
By Mohamad Amin Mirzaei, Pantea Amoie, Ali Ekhterachian, Matin Mirzababaei, Babak Khalaj
arXiv:2606. 19733v1 Announce Type: cross Abstract: Efficiently retrieving specific 3D instances from large-scale scenes via natural language prompts remains a formidable challenge in multimedia analysis.
By Xiuyuan Zhu, Ke Lu, Zijie Yang, Chao Yue, Jian Xue, Dongming Zhang
arXiv:2507.22052v3 Announce Type: replace
Abstract: We present Ov3R, a novel framework for open-vocabulary semantic 3D reconstruction from RGB video streams, designed to advance Spatial AI. The syste...
By Ziren Gong, Xiaohan Li, Fabio Tosi, Jiawei Han, Stefano Mattoccia, Jianfei Cai, Matteo Poggi
Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstructed environments using natural language. However, current 3D Gaussian frameworks are often bottlenecked by restrictive multiview capture requirements, costly scene-specific optimization, and the massive memory overhead of storing dense language features.