Hugging Face Trending Papers

PC-Edit: Prompt-Contrastive Region Discovery and Region-Guided Editing

Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by the source silhouette, and preservation of unrelated content. Existing training-free editors either localize edits from terminal predictions under source and target prompts or preserve unrelated content through spatially unselective source-feature reuse without explicit region discovery.

arXiv Computer Vision
Sep 18

Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network

RefineEdit is a training‑free prompt‑to‑prompt image editing framework that uses a Generative Refinement Network to edit images by refining binary image codes. It couples edit localization with content generation, selecting editable positions based on signed probability differences between an editing branch and a source branch, and stabilizes edits with adaptive spatial freezing and finite bit locking. The method requires no additional training, external masks, or attention control, and outperforms other methods on PIE‑Bench in background‑preservation metrics and CLIP scores.

By Yulong Chen, Ziqian Zhang, Haoyu Zhang, Ao He, Senmao Li, Kai Wang
arXiv AI
1d ago

VibeEdit: Image Editing with Canvas Instructions

VibeEdit is a new image‑editing tool that lets users draw spatial marks and add brief notes directly on an image to create a canvas instruction. These instructions guide the model to perform tasks such as adding, removing, replacing, or moving objects, as well as modifying attributes, without needing a separate text prompt. The system is trained on 1.55 million annotated edit pairs and uses a layer‑decoupled conditioning approach, achieving higher VLM rubric scores and PSNR on a benchmark that tests target selection among similar objects.

By Jinjing Zhao, Fangyun Wei, Yitong Wang, Xiuyu Wu, Yunuo Chen, Yang Yue, Sirui Zhang, Wenbo Wang, Hongyang Zhang, Dong Chen, Yan Lu, Chang Xu
arXiv Computer Vision
Sep 3

SR-Edit: Region-Aware Image Editing via Self-Refinement

SR-Edit is a new image editing framework that uses iterative self‑refinement to improve fidelity. At each step it extracts precise, self‑consistent region separations from the model’s predictions and then enforces preservation in non‑edit areas with correction updates that stay aligned with the original sampling dynamics. Experiments show that SR‑Edit delivers better preservation and overall image quality than existing editing techniques.

By Andong Wang, Zehua Chen, Yuxuan Jiang, Jun Zhu
Hugging Face Trending Papers
Aug 18

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

CoinVE-200K is a large, high‑quality dataset for compositional instruction‑guided video editing, featuring 1080p video‑editing pairs up to 201 frames long and containing 2–5 atomic editing operations per sample. The dataset covers diverse editing intents—targeting humans, objects, and backgrounds with addition, removal, modification, and stylization—while ensuring instruction faithfulness, visual quality, temporal consistency, and compositional diversity through a careful generation and filtering pipeline. CoinVE-Bench benchmarks these capabilities, and CoinVE-Edit, a 22B model built on Wan2.1‑T2V‑14B and Qwen3‑VL‑8B‑Instruct, demonstrates strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.

arXiv AI
Jun 15

HiLo-Token: Input-Adaptive High-Low Frequency Token Compression for Efficient Image Editing

arXiv:2606. 13898v1 Announce Type: cross Abstract: Creative image editing tools, such as Photoshop's Remove or Generative Fill buttons, are central to everyday customer use and account for a major share of traffic in Photoshop and Lightroom.

By Haoran You, Yotam Nitzan, Lingzhi Zhang, Yifan Gong, Mang-Tik Chiu, Connelly Barnes, Yan Kang, Yuqian Zhou, Eli Shechtman, Sohrab Amirghodsi
arXiv Computer Vision
Sep 21

Edit-VAR: Taming Visual Autoregressive Model for Precise Video Editing

Edit‑VAR is a training‑free, inversion‑free framework that uses a pretrained visual autoregressive video model for text‑guided video editing. It encodes the source video into multi‑scale discrete tokens and applies probability‑guided conditional token replacement, attention‑guided token‑wise and scale‑aware modulation, and scale‑decoupled generation to preserve source appearance while enabling precise edits. The method also includes residual‑guided token pruning to reduce inference cost, and experimental results show it outperforms existing training‑free video editing methods in fidelity, source preservation, temporal coherence, and efficiency.

By Chongbo Zhao, Jiangming Wang, Xilai Wang, Xinyu Wang, Jingyi Tang, Chunjie Hao, Pengjie Song, Yue Ma
arXiv Computer Vision
Sep 11

Overpainting: Localized Context-aware Diffusion Image Editing

The paper introduces "overpainting," a localized, context-aware image editing technique that allows users to specify precise or loose editing regions via a trimap. The method adapts a pretrained diffusion model with joint attention and low‑rank adaptation, incorporating attention‑dropout to balance noise, source, and mask inputs. An automated pipeline generates training data by pairing images from language‑based editing models, curating them, and extracting trimaps, enabling the model to perform a wide range of editing tasks.

By Sam Sartor, Iliyan Georgiev, Michael Fischer, Valentin Deschaintre, Pieter Peers
arXiv Computer Vision
Aug 25

Reference-free Human-Object Interaction Editing

arXiv:2503.09130v2 Announce Type: replace-cross Abstract: This paper presents InteractEdit, a novel framework for reference-free Human-Object Interaction (HOI) editing that tackles the challenging ta...

By Jiun Tian Hoe, Weipeng Hu, Wei Zhou, Chao Xie, Ziwei Wang, Xudong Jiang, Yap-Peng Tan, Chee Seng Chan