EditaLive! is a new real‑time framework for character video editing in live streaming, built on a pretrained image animation model (Wan‑Animate) that separates appearance from motion. It uses the CharEdit‑50K dataset for reference‑frame editing and video reconstruction, and adapts the model from offline bidirectional to causal streaming generation. A self‑rollout distillation strategy compresses the model into a two‑step sampler, employing fixed RoPE, alignment forcing, and first‑frame preserved sparse attention to reduce appearance drift and achieve low‑latency inference while preserving facial expressions.
By Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang, Bo Li, Xiaodong Cun
Edit‑VAR is a training‑free, inversion‑free framework that uses a pretrained visual autoregressive video model for text‑guided video editing. It encodes the source video into multi‑scale discrete tokens and applies probability‑guided conditional token replacement, attention‑guided token‑wise and scale‑aware modulation, and scale‑decoupled generation to preserve source appearance while enabling precise edits. The method also includes residual‑guided token pruning to reduce inference cost, and experimental results show it outperforms existing training‑free video editing methods in fidelity, source preservation, temporal coherence, and efficiency.
By Chongbo Zhao, Jiangming Wang, Xilai Wang, Xinyu Wang, Jingyi Tang, Chunjie Hao, Pengjie Song, Yue Ma
arXiv:2607. 19895v1 Announce Type: cross Abstract: Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion.
By Habin Lim, Gyeong-Moon Park
Recent diffusion editors perform diverse instruction-based edits while conditioning on the source image at every denoising step. Yet persistent source-image conditioning can limit how fully an edit is executed and how natural the result appears, especially when the target scene diverges substantially from the input.
arXiv:2606. 19676v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success in image and video generation and editing.
By Haengbok Chung
arXiv:2606. 01014v1 Announce Type: cross Abstract: We address text-based 3D human motion editing, where the goal is to preserve the style and structure of a source motion while applying edits described in natural language.
By Gyojin Han, Junmo Kim
arXiv:2609.40356v1 Announce Type: cross
Abstract: Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits th...
By Xinghao Chen, Xiangbo Gao, Jiongze Yu, Yuheng Wu, Zhengzhong Tu
RefineEdit is a training‑free prompt‑to‑prompt image editing framework that uses a Generative Refinement Network to edit images by refining binary image codes. It couples edit localization with content generation, selecting editable positions based on signed probability differences between an editing branch and a source branch, and stabilizes edits with adaptive spatial freezing and finite bit locking. The method requires no additional training, external masks, or attention control, and outperforms other methods on PIE‑Bench in background‑preservation metrics and CLIP scores.
By Yulong Chen, Ziqian Zhang, Haoyu Zhang, Ao He, Senmao Li, Kai Wang
arXiv:2609.36832v1 Announce Type: new
Abstract: Text-to-video (T2V) diffusion models can generate realistic depictions of actions such as kicking, stabbing, and shooting, raising safety concerns that...
By Ping Liu, Chi Zhang
arXiv:2607.18227v2 Announce Type: replace
Abstract: In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and imag...
By Dingyun Zhang, Lixue Gong, Wei Liu
arXiv:2609.33167v2 Announce Type: replace
Abstract: We present FloodDiffusion 2 (FD2), an efficient and controllable framework that builds upon FloodDiffusion (FD1), a state-of-the-art streaming moti...
By Yiyi Cai, Yuhan Wu, Kunhang Li, Tu Fangyuan, Xiangyue Zhang, Qiaoge Li, Zhixiang Wang, Kaipeng Zhang, Haiyang Liu
ContextAnyone is a context‑aware diffusion framework that treats a reference image as an explicitly preserved appearance anchor rather than a simple conditioning signal. By jointly reconstructing the reference image and generating the target video within a shared diffusion transformer, it provides direct supervision for maintaining identity and fine‑grained appearance throughout denoising. The method introduces asymmetric information flow and Gap‑RoPE positional representations to keep the reference stable while allowing selective access by video tokens, and demonstrates improved identity and appearance consistency on an OpenVid‑HD benchmark.
By Ziyang Mai, Yu-Wing Tai