This paper presents a zero‑shot framework for removing objects from real images using a frozen Stable Diffusion model, avoiding any task‑specific training. The pipeline combines SAM‑based mask construction, BLIP caption conditioning, DDIM inversion, background‑weighted masked null‑text optimization, decoder self‑attention masking, hard outside‑mask latent anchoring, and localized renoise‑denoise refinement. Experiments show effective removal of objects and context‑consistent replacement, with background‑weighted NTI especially helpful for complex backgrounds and repeated refinement reducing residual artifacts.
By Arman Taghizadeh (Institute of Cognitive Science, Osnabr\"uck University, Osnabr\"uck, Germany), Ulf Krumnack (Institute of Cognitive Science, Osnabr\"uck University, Osnabr\"uck, Germany), Kai-Uwe K\"uhnberger (Institute of Cognitive Science, Osnabr\"uck University, Osnabr\"uck, Germany)
Removing an object from a real image requires more than synthesizing plausible content within a mask: the method must suppress residual object features, preserve the unedited scene, and generate repla...
arXiv:2609.21424v1 Announce Type: new
Abstract: Few-shot semantic segmentation (FSS) of strip steel surface defects (S$^3$D) has posed significant challenges distinct from natural scenes. Unlike natu...
By Qian Xu, Hang Xiong, Anpeng Wang, Sam Kwong, Cong Zhang, Runmin Cong
The paper introduces a cross‑modal pseudo‑labeling pipeline for unsupervised domain adaptation in semantic segmentation, particularly for waste sorting. It combines SAM for class‑agnostic region proposals with EVA‑CLIP to assign semantic labels via region‑text similarity, applying confidence filtering to ensure reliable pseudo‑labels for self‑training. An optional BLIP‑based language‑grounded verification further refines ambiguous regions, and the method shows consistent improvements over source‑only baselines on synthetic‑to‑real driving and lab‑to‑factory waste sorting shifts.
By Udo Schlegel, Shubhangi, Gabriel Dax, Sai Rahul Kaminwar, Florian Karl, Thomas Seidl
arXiv:2607. 02404v1 Announce Type: cross Abstract: Image encoders trained with LeJEPA can deliver strong features for downstream tasks, but, like other image-level self-supervised methods, typically require large training datasets.
By Jakob Geusen, Ender Konukoglu
arXiv:2609.25500v1 Announce Type: new
Abstract: Training data quantity and quality greatly affect object detection model performance, regardless of model architecture. When using object detection mod...
By Lonny Lundsten, Kevin Barnard, Dave Caress
SR‑Ground is a large‑scale dataset created to enable fine‑grained segmentation of visual artifacts in super‑resolved images. It contains 63,000 images processed by various state‑of‑the‑art SR models, each annotated at the pixel level for six distinct artifact types, validated through a crowdsourcing study with 1,062 participants. The dataset improves the training of image quality assessment models with grounding capabilities and supports a fine‑tuning pipeline that reduces perceptible artifacts in SR outputs, outperforming no‑reference methods on both benchmark and real‑world low‑resolution datasets.
By Artem Borisov, Evgeney Bogatyrev, Khaled Abud, Dmitriy Vatolin
arXiv:2610.01022v1 Announce Type: new
Abstract: Memory-attention-based Video Instance Segmentation (VIS) methods have demonstrated strong zero-shot tracking capability, yet their substantial memory r...
By Arash Rocky, Q. M. Jonathan Wu
arXiv:2608.29917v1 Announce Type: new
Abstract: Personalized segmentation and personalized retrieval both aim to identify the same physical object across different images. While the former localizes...
By Gabriele Trivigno, Marcos Alfaro, Claudia Cuttano, Gabriele Berton, Luis Pay\'a, Carlo Masone
arXiv:2608.22950v1 Announce Type: new
Abstract: Floating waste in inland waterways threatens aquatic ecosystems and requires timely monitoring under cluttered, multi-object conditions. Existing aquat...
By Md. Asaduzzaman Shuvo, Ahsan Farabi, Md. Abdul Ahad Minhaz, Mahedi Hasan, Israt Khandaker, Ibrahim Khalil Shanto, Muhammad Nomani Kabir
arXiv:2606. 15786v1 Announce Type: cross Abstract: The advent of large pretrained foundation models for computer vision has significantly improved the efficiency of visual data interpretation.
By Aniq Ahmad, Heather Bedle, Ahmad Mustafa
Floating waste in inland waterways threatens aquatic ecosystems and requires timely monitoring under cluttered, multi-object conditions. Existing aquatic-waste datasets provide limited geographic cove...