arXiv AI By Sung-Hoon Yoon, Hoyong Kwon, Changgyoon Oh, Kuk-Jin Yoon

DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation

Read the original on arXiv AI →

arXiv:2607. 21371v1 Announce Type: cross Abstract: Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 15

ViCo-SAM3: Vision-Conditioned Alignment for Open-Vocabulary Camouflaged Object Segmentation

ViCo-SAM3 introduces a Vision-Conditioned alignment framework for open-vocabulary camouflaged object segmentation. The approach adds a vision-conditioned (ViCo) module that dynamically adjusts text embeddings based on global visual context, and a vision-conditioned cross-modal binding (ViCoBind) module to improve interaction between visual and textual representations. These innovations close the semantic gap between text and pixel-level cues, enabling state‑of‑the‑art performance on the OVCamo benchmark without heavy parameter overhead.

By Qiangqiang Zhou, Wenjun Tang, Yong Chen, Dandan Zhu, Jiawei Xu
arXiv AI
Sep 10

GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

GoDeep is an annotation‑free method for open‑vocabulary 3D scene understanding that uses a vision‑language model solely as a translator to generate structured, entity‑level descriptions of each image. These descriptions are projected and aggregated in a language‑only embedding space, eliminating the need for a 3D training corpus or domain‑specific encoder. The approach achieves competitive performance on ScanNet++ and a cultural heritage benchmark, accurately localizes out‑of‑vocabulary objects, and offers explainable, point‑level predictions.

By Thodoris Betsas, Anastasios Doulamis, Andreas Georgopoulos