The paper introduces Text-to-Seed (T2S), a training‑free framework for open‑vocabulary semantic segmentation that repurposes Stable Diffusion to generate attention‑based seed points from text queries. These sparse seeds serve as point prompts for the Segment Anything Model (SAM), enabling reliable region expansion without relying on inaccurate coarse masks. T2S achieves strong performance on standard OVSS benchmarks using only the text‑to‑region correspondence of diffusion models and no task‑specific training or extra annotations.
By Kumju Jo, Heesun Jung, Sungyong Baik
arXiv:2606. 31603v1 Announce Type: cross Abstract: Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.
By Nikolai R\"ohrich, Julian Glei{\ss}ner, Ahmed H. A. Ibrahim, Silvan Mertes, Tobias Huber
arXiv:2608.29917v1 Announce Type: new
Abstract: Personalized segmentation and personalized retrieval both aim to identify the same physical object across different images. While the former localizes...
By Gabriele Trivigno, Marcos Alfaro, Claudia Cuttano, Gabriele Berton, Luis Pay\'a, Carlo Masone
arXiv:2607. 13421v1 Announce Type: cross Abstract: Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural language expression.
By Kai Chen, Ming Dai, Wenxuan Cheng, Wankou Yang
Personalized segmentation and personalized retrieval both aim to identify the same physical object across different images. While the former localizes the object within a target image, the latter retr...
arXiv:2607. 26107v1 Announce Type: cross Abstract: Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions.
By Xinran Liu, Shouqian Shi, Yutong Chen, Ge Wang, Xin-Wei Yao, Sheng Zhong
SOCO is a new benchmark for Semantic Object Correspondence that introduces a taxonomy of correspondence types and provides consistent, functionally meaningful keypoint annotations across 100 categories and over 1M correspondence pairs. It also includes keypoint language descriptions, enabling evaluation of large vision‑language models and their fine‑grained part‑level understanding. Experiments show that vision foundation backbones encode strong semantic structure but transfer correspondences poorly across related categories, LVLMs excel at text‑prompted part localization but lag in visual‑reference matching, and correspondence performance predicts dense downstream tasks more strongly than ImageNet classification.
By Olaf D\"unkel, Basavaraj Sunagad, Haoran Wang, David T. Hoffmann, Christian Theobalt, Adam Kortylewski
arXiv:2502. 06818v4 Announce Type: replace Abstract: Recent works modify CLIP to perform open-vocabulary semantic segmentation in a training-free manner (TF-OVSS).
By Jingyun Wang, Cilin Yan, Guoliang Kang
GazeRefine is a training‑free framework that uses eye‑gaze data as an inference‑time prompt for zero‑shot medical image segmentation. It converts sparse, duration‑weighted fixations into foreground and background priors that initialize semantic prototypes in a frozen DINOv3 feature space, then iteratively refines these prototypes through discrimination, affinity propagation, and anchoring to the gaze guidance. The method achieves strong results on colonoscopy polyp segmentation and competitive performance on prostate MRI, demonstrating that gaze‑guided prototype refinement can enable segmentation without dense expert annotations or model fine‑tuning.
By Mohammed Oussama Benyahia, Marouane Tliba, Mohamed Amine Kerkouri, Taifour Yousra, Bin Wang, Max Bengtsson, Gorkem Durak, Elif Keles, Zuheng Ming, Marek Penhaker, Azeddine Beghdadi, Ulas Bagci, Aladine Chetouani
arXiv:2501.04001v4 Announce Type: replace
Abstract: This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-...
By Haobo Yuan, Xiangtai Li, Tao Zhang, Yueyi Sun, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang
arXiv:2509. 24528v4 Announce Type: replace-cross Abstract: Object retrieval from a scene has become a new trend of research due to its numerous applications.
By Mohamad Amin Mirzaei, Pantea Amoie, Ali Ekhterachian, Matin Mirzababaei, Babak Khalaj
arXiv:2503. 09399v4 Announce Type: replace-cross Abstract: Large-scale image classification datasets exhibit strong compositional biases: objects tend to be centered, appear at characteristic scales, and co-occur with class-specific context.
By Tobias Christian Nauen, Brian Moser, Federico Raue, Stanislav Frolov, Andreas Dengel