arXiv AI

DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation

arXiv:2607. 21371v1 Announce Type: cross Abstract: Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories.

arXiv Computer Vision
Sep 15

ViCo-SAM3: Vision-Conditioned Alignment for Open-Vocabulary Camouflaged Object Segmentation

ViCo-SAM3 introduces a Vision-Conditioned alignment framework for open-vocabulary camouflaged object segmentation. The approach adds a vision-conditioned (ViCo) module that dynamically adjusts text embeddings based on global visual context, and a vision-conditioned cross-modal binding (ViCoBind) module to improve interaction between visual and textual representations. These innovations close the semantic gap between text and pixel-level cues, enabling state‑of‑the‑art performance on the OVCamo benchmark without heavy parameter overhead.

By Qiangqiang Zhou, Wenjun Tang, Yong Chen, Dandan Zhu, Jiawei Xu
arXiv AI
Sep 10

GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

GoDeep is an annotation‑free method for open‑vocabulary 3D scene understanding that uses a vision‑language model solely as a translator to generate structured, entity‑level descriptions of each image. These descriptions are projected and aggregated in a language‑only embedding space, eliminating the need for a 3D training corpus or domain‑specific encoder. The approach achieves competitive performance on ScanNet++ and a cultural heritage benchmark, accurately localizes out‑of‑vocabulary objects, and offers explainable, point‑level predictions.

By Thodoris Betsas, Anastasios Doulamis, Andreas Georgopoulos
arXiv Computer Vision
6d ago

TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking

TRACKGRAPH is an online open‑vocabulary 3D mapping system that tracks 2D masks in the image stream before fusing them into a class‑agnostic 3D segment layer within a hierarchical scene graph. It uses FastSAM and CLIP for sparse keyframes, DINOv3 for dense mask propagation, and compact multi‑view CLIP embeddings for open‑vocabulary retrieval. The method outperforms state‑of‑the‑art mapping techniques on Replica, ScanNet++, and HM3D, achieving higher synonym frequency, faster processing, and lower GPU memory usage, and has been deployed on quadruped robots and drones at real‑time rates.

By Peder Borge Hellesylt, Albert Gassol Puigjaner, Kostas Alexis, Annette Stahl
arXiv AI
Jun 26

Bridging Vision and Language Concepts through Optimal Transport Semantic Flow

arXiv:2606. 26891v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual and textual representations are aligned or matched.

By Chenyang Zhang, Anqi Dong, Guangming Zhu, Nuoye Xiong, Siyuan Wang, Lin Mei, Liang Zhang
arXiv Computer Vision
Aug 28

DINOcular: Self-Supervised Visuospatial Representations

DINOcular is a self‑supervised framework that learns joint visuospatial representations from RGB‑D observations. It fuses depth‑derived geometric priors with a visual backbone using inter‑patch and intra‑patch fusion, allowing the model to encode both appearance and spatial structure efficiently. The resulting representation improves 3D awareness on multiple geometry benchmarks while staying competitive on standard RGB‑D semantic segmentation tasks.

By Farkhat Almukhamedov, Sami Azirar, Hermann Blum