arXiv Computer Vision By Qiangqiang Zhou, Wenjun Tang, Yong Chen, Dandan Zhu, Jiawei Xu

ViCo-SAM3: Vision-Conditioned Alignment for Open-Vocabulary Camouflaged Object Segmentation

Read the original on arXiv Computer Vision →

ViCo-SAM3 introduces a Vision-Conditioned alignment framework for open-vocabulary camouflaged object segmentation. The approach adds a vision-conditioned (ViCo) module that dynamically adjusts text embeddings based on global visual context, and a vision-conditioned cross-modal binding (ViCoBind) module to improve interaction between visual and textual representations. These innovations close the semantic gap between text and pixel-level cues, enabling state‑of‑the‑art performance on the OVCamo benchmark without heavy parameter overhead.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Jul 15

Fine-grained CLIP fine-tuning with self-annotated region alignment

Contrastive Language-Image Pre-training (CLIP) has been shown to have limitations in its fine-grained dense feature representation, due to its pre-training focusing on matching the whole image to a text description. Considering the large data and computational burden in pre-training a vision-language model from scratch, a series of works aim to enhance the fine-grained ability of CLIP through a fine-tuning scheme.