Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories. While the self-supervised model DINOv3 provides strong structured visual representations, its lack of native textual alignment hinders its direct application to OVSS.
arXiv:2608. 03023v1 Announce Type: cross Abstract: Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods.
By Changhao Zhao, Haoxiang Li, Yuke Li, Hai Liu, LingLin Zeng
arXiv:2609.28187v1 Announce Type: new
Abstract: Self-supervised vision transformers trained with DINO-style objectives exhibit striking emergent semantic representation quality across visual tasks, y...
By Basavaraj Sunagad, Artur Jesslen, Adam Kortylewski
ViCo-SAM3 introduces a Vision-Conditioned alignment framework for open-vocabulary camouflaged object segmentation. The approach adds a vision-conditioned (ViCo) module that dynamically adjusts text embeddings based on global visual context, and a vision-conditioned cross-modal binding (ViCoBind) module to improve interaction between visual and textual representations. These innovations close the semantic gap between text and pixel-level cues, enabling state‑of‑the‑art performance on the OVCamo benchmark without heavy parameter overhead.
By Qiangqiang Zhou, Wenjun Tang, Yong Chen, Dandan Zhu, Jiawei Xu
arXiv:2609.24564v1 Announce Type: new
Abstract: CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encode...
By Zelin Peng, Zhengqin Xu, Changsong Wen, Yu Huang, Yaoming Wang, Xiaokang Yang, Wei Shen
GoDeep is an annotation‑free method for open‑vocabulary 3D scene understanding that uses a vision‑language model solely as a translator to generate structured, entity‑level descriptions of each image. These descriptions are projected and aggregated in a language‑only embedding space, eliminating the need for a 3D training corpus or domain‑specific encoder. The approach achieves competitive performance on ScanNet++ and a cultural heritage benchmark, accurately localizes out‑of‑vocabulary objects, and offers explainable, point‑level predictions.
By Thodoris Betsas, Anastasios Doulamis, Andreas Georgopoulos
TRACKGRAPH is an online open‑vocabulary 3D mapping system that tracks 2D masks in the image stream before fusing them into a class‑agnostic 3D segment layer within a hierarchical scene graph. It uses FastSAM and CLIP for sparse keyframes, DINOv3 for dense mask propagation, and compact multi‑view CLIP embeddings for open‑vocabulary retrieval. The method outperforms state‑of‑the‑art mapping techniques on Replica, ScanNet++, and HM3D, achieving higher synonym frequency, faster processing, and lower GPU memory usage, and has been deployed on quadruped robots and drones at real‑time rates.
By Peder Borge Hellesylt, Albert Gassol Puigjaner, Kostas Alexis, Annette Stahl
arXiv:2606. 26891v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual and textual representations are aligned or matched.
By Chenyang Zhang, Anqi Dong, Guangming Zhu, Nuoye Xiong, Siyuan Wang, Lin Mei, Liang Zhang
arXiv:2607. 08541v1 Announce Type: cross Abstract: Open-vocabulary object detection and segmentation aim to recognize arbitrary objects beyond predefined categories.
By ZhiXin Sun
arXiv:2608.21136v1 Announce Type: new
Abstract: Recently, open-vocabulary zero-shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data-intensive supe...
By Jie Xu, Na Zhao
DINOcular is a self‑supervised framework that learns joint visuospatial representations from RGB‑D observations. It fuses depth‑derived geometric priors with a visual backbone using inter‑patch and intra‑patch fusion, allowing the model to encode both appearance and spatial structure efficiently. The resulting representation improves 3D awareness on multiple geometry benchmarks while staying competitive on standard RGB‑D semantic segmentation tasks.
By Farkhat Almukhamedov, Sami Azirar, Hermann Blum
arXiv:2607. 00858v1 Announce Type: cross Abstract: Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled representations.
By Peiyuan Zhu, Shaoan Xie, Zijian Li, Yifan Shen, Namrata Deka, Harsh Shrivastava, Guangyi Chen, Kun Zhang