arXiv AI By Xiuyuan Zhu, Ke Lu, Zijie Yang, Chao Yue, Jian Xue, Dongming Zhang

QueryGaussian: Scalable and Training-Free Open-Vocabulary 3D Instance Retrieval

Read the original on arXiv AI →

arXiv:2606. 19733v1 Announce Type: cross Abstract: Efficiently retrieving specific 3D instances from large-scale scenes via natural language prompts remains a formidable challenge in multimedia analysis.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 30

MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians

Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstructed environments using natural language. However, current 3D Gaussian frameworks are often bottlenecked by restrictive multiview capture requirements, costly scene-specific optimization, and the massive memory overhead of storing dense language features.

arXiv AI
Aug 25

ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian Splatting

ExtrinSplat is a new framework that separates geometry from semantics in 3D Gaussian Splatting scenes. It clusters Gaussians into overlapping 3D object groups and uses a Vision‑Language Model to generate lightweight textual hypotheses, creating an extrinsic index layer that handles complex polysemy. This approach reduces adaptation time from hours to minutes, cuts storage overhead by orders of magnitude, and outperforms existing embedding‑based methods on open‑vocabulary 3D object selection and semantic segmentation benchmarks.

By Jiayu Ding, Xinpeng Liu, Zhiyi Pan, Shiqiang Long, Ge Li
arXiv Computer Vision
Sep 18

Scene-Q: Confidence-Aware Coarse-to-Fine Querying of 3D Scenes with Selective VLM Reasoning

Scene-Q is a confidence‑aware, coarse‑to‑fine querying framework for indoor mobile robots that grounds natural‑language queries in a 3D map. It normalizes encoder scores with temperature scaling and only invokes a reasoning VLM for low‑confidence cases, while high‑confidence queries are answered by fast retrieval. The method improves open‑vocabulary 3D instance segmentation on ScanNet200 and natural‑language 3D instance retrieval on real‑world reconstructions, especially for spatial and relational queries, while maintaining a substantial fraction of queries on the fast path.

By Juno Kim, Yesol Park, Hye-Jung Yoon, Byoung-Tak Zhang