This paper presents a zero‑shot framework for removing objects from real images using a frozen Stable Diffusion model, avoiding any task‑specific training. The pipeline combines SAM‑based mask construction, BLIP caption conditioning, DDIM inversion, background‑weighted masked null‑text optimization, decoder self‑attention masking, hard outside‑mask latent anchoring, and localized renoise‑denoise refinement. Experiments show effective removal of objects and context‑consistent replacement, with background‑weighted NTI especially helpful for complex backgrounds and repeated refinement reducing residual artifacts.
By Arman Taghizadeh (Institute of Cognitive Science, Osnabr\"uck University, Osnabr\"uck, Germany), Ulf Krumnack (Institute of Cognitive Science, Osnabr\"uck University, Osnabr\"uck, Germany), Kai-Uwe K\"uhnberger (Institute of Cognitive Science, Osnabr\"uck University, Osnabr\"uck, Germany)
Removing an object from a real image requires more than synthesizing plausible content within a mask: the method must suppress residual object features, preserve the unedited scene, and generate repla...
arXiv:2609.21424v1 Announce Type: new
Abstract: Few-shot semantic segmentation (FSS) of strip steel surface defects (S$^3$D) has posed significant challenges distinct from natural scenes. Unlike natu...
By Qian Xu, Hang Xiong, Anpeng Wang, Sam Kwong, Cong Zhang, Runmin Cong
The paper introduces a cross‑modal pseudo‑labeling pipeline for unsupervised domain adaptation in semantic segmentation, particularly for waste sorting. It combines SAM for class‑agnostic region proposals with EVA‑CLIP to assign semantic labels via region‑text similarity, applying confidence filtering to ensure reliable pseudo‑labels for self‑training. An optional BLIP‑based language‑grounded verification further refines ambiguous regions, and the method shows consistent improvements over source‑only baselines on synthetic‑to‑real driving and lab‑to‑factory waste sorting shifts.
By Udo Schlegel, Shubhangi, Gabriel Dax, Sai Rahul Kaminwar, Florian Karl, Thomas Seidl
arXiv:2607. 02404v1 Announce Type: cross Abstract: Image encoders trained with LeJEPA can deliver strong features for downstream tasks, but, like other image-level self-supervised methods, typically require large training datasets.
By Jakob Geusen, Ender Konukoglu
arXiv:2609.25500v1 Announce Type: new
Abstract: Training data quantity and quality greatly affect object detection model performance, regardless of model architecture. When using object detection mod...
By Lonny Lundsten, Kevin Barnard, Dave Caress