arXiv AI

VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval

arXiv:2607. 08541v1 Announce Type: cross Abstract: Open-vocabulary object detection and segmentation aim to recognize arbitrary objects beyond predefined categories.

arXiv AI
Aug 19

Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing

The paper introduces OVRSISBench, a unified benchmark for open‑vocabulary remote sensing image segmentation, and evaluates existing OVS/OVRSIS models, uncovering their shortcomings in remote sensing contexts. Leveraging insights from this evaluation, the authors propose RSKT‑Seg, a new framework featuring a Multi‑Directional Cost Map Aggregation module, an Efficient Cost Map Fusion transformer, and a Remote Sensing Knowledge Transfer module. Experiments on the benchmark demonstrate that RSKT‑Seg outperforms strong baselines by +3.8 mIoU and +5.9 mACC while achieving twice the inference speed.

By Bingyu Li, Haocheng Dong, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li
arXiv Computer Vision
6d ago

TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking

TRACKGRAPH is an online open‑vocabulary 3D mapping system that tracks 2D masks in the image stream before fusing them into a class‑agnostic 3D segment layer within a hierarchical scene graph. It uses FastSAM and CLIP for sparse keyframes, DINOv3 for dense mask propagation, and compact multi‑view CLIP embeddings for open‑vocabulary retrieval. The method outperforms state‑of‑the‑art mapping techniques on Replica, ScanNet++, and HM3D, achieving higher synonym frequency, faster processing, and lower GPU memory usage, and has been deployed on quadruped robots and drones at real‑time rates.

By Peder Borge Hellesylt, Albert Gassol Puigjaner, Kostas Alexis, Annette Stahl
arXiv AI
Aug 25

ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian Splatting

ExtrinSplat is a new framework that separates geometry from semantics in 3D Gaussian Splatting scenes. It clusters Gaussians into overlapping 3D object groups and uses a Vision‑Language Model to generate lightweight textual hypotheses, creating an extrinsic index layer that handles complex polysemy. This approach reduces adaptation time from hours to minutes, cuts storage overhead by orders of magnitude, and outperforms existing embedding‑based methods on open‑vocabulary 3D object selection and semantic segmentation benchmarks.

By Jiayu Ding, Xinpeng Liu, Zhiyi Pan, Shiqiang Long, Ge Li