Zero-shot image segmentation with CLIPSeg
Related stories
Universal Image Segmentation with Mask2Former and OneFormer
MARS-CLIP: Multi-Resolution and Attention Refined Zero-Shot Image Segmentation
MARS-CLIP is a zero‑shot semantic segmentation framework that builds on CLIP by adding a multi‑resolution feature extraction module and an attention refinement mechanism. The multi‑resolution module fuses fine‑grained local features with global context to mitigate low spatial resolution, while the attention refinement injects spatial and color biases from intermediate layers into the final self‑attention block to better recover object boundaries. Experiments on six public datasets show that MARS‑CLIP outperforms state‑of‑the‑art methods.
GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation
GazeRefine is a training‑free framework that uses eye‑gaze data as an inference‑time prompt for zero‑shot medical image segmentation. It converts sparse, duration‑weighted fixations into foreground and background priors that initialize semantic prototypes in a frozen DINOv3 feature space, then iteratively refines these prototypes through discrimination, affinity propagation, and anchoring to the gaze guidance. The method achieves strong results on colonoscopy polyp segmentation and competitive performance on prostate MRI, demonstrating that gaze‑guided prototype refinement can enable segmentation without dense expert annotations or model fine‑tuning.
Scene-Centric Unsupervised Video Panoptic Segmentation
Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any human supervision.
Seeing as Humans Do: Learning from Motion to Segment Anything Without Supervision
arXiv:2609.39785v1 Announce Type: new Abstract: The Segment Anything Model (SAM) relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling. While unsupervised...
Don't waste SAM
Meta AI has recently released the Segment Anything Model (SAM), which demonstrates exceptional zero-shot image segmentation performance across various tasks with remarkable accuracy. Despite its inability to provide accurate segmentation across multiple research fields, SAM still serves as a valuable starting point for supporting the segmentation pipeline process, particularly for tasks that require extensive and senior skills annotations.
Zero-Shot Object Removal via Attention Masking, Latent Anchoring, and Refinement
Removing an object from a real image requires more than synthesizing plausible content within a mask: the method must suppress residual object features, preserve the unedited scene, and generate repla...
Diffusion Editing with Soft Mask: Pixel Level Redo of Image and Video with Adjustable Strength
Diffusion Editing with Soft Mask: Pixel Level Redo of Image and Video with Adjustable Strength introduces SoftPaint, a zero‑shot sampling method that uses soft masks to provide continuous, pixel‑level control over edits in diffusion models. The approach employs a Langevin‑iteration sampler that respects per‑pixel mask strengths, enabling smooth edits from preserving to fully re‑synthesizing content across image and video backbones. SoftPaint is gradient‑free, memory‑efficient, and works universally with existing diffusion models.
Zero-Shot Object Removal via Attention Masking, Latent Anchoring, and Refinement
This paper presents a zero‑shot framework for removing objects from real images using a frozen Stable Diffusion model, avoiding any task‑specific training. The pipeline combines SAM‑based mask construction, BLIP caption conditioning, DDIM inversion, background‑weighted masked null‑text optimization, decoder self‑attention masking, hard outside‑mask latent anchoring, and localized renoise‑denoise refinement. Experiments show effective removal of objects and context‑consistent replacement, with background‑weighted NTI especially helpful for complex backgrounds and repeated refinement reducing residual artifacts.
Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets
arXiv:2606. 07032v1 Announce Type: cross Abstract: Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve a target image based on a query composed of a reference image and a relative caption without training samples.
Mask Proposal Voting Based on Geodesic Framework for Robust Image Segmentation
arXiv:2606. 14912v1 Announce Type: cross Abstract: Despite great advances, finding accurate segmentation remains a challenging task, especially in scenarios with cluttered backgrounds, complex intensity variations and topology appearance.