Memory-V2V is a memory‑augmented video‑to‑video diffusion framework designed to improve cross‑turn consistency in multi‑turn video editing. It stores previous outputs in an external memory, retrieves relevant edits, and incorporates them via relevance‑aware tokenization and adaptive compression, allowing scalable conditioning without linear computational growth. Experiments on iterative video novel view synthesis and text‑guided long video editing show that Memory‑V2V enhances consistency while preserving visual quality and outperforming strong baselines with modest overhead.
By Dohun Lee, Chun-Hao Paul Huang, Xuelin Chen, Jong Chul Ye, Duygu Ceylan, Hyeonho Jeong
arXiv:2606. 05950v1 Announce Type: new Abstract: Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models.
By Yuxiao Ye, Haoran He, Fangyuan Kong, Xintao Wang, Pengfei Wan, Kun Gai, Ling Pan
arXiv:2607. 19895v1 Announce Type: cross Abstract: Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion.
By Habin Lim, Gyeong-Moon Park
Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday users. However, editing models trained primarily for single-turn edits often break down in multi-turn editing--the natural interactive setting where a user iteratively refines an image based on the model's own previous outputs.
The paper introduces EditVid, a training‑free framework for video editing that unifies instruction‑guided and reference‑guided edits. It employs sparse causal memory for local coherence, correspondence‑based post‑attention token injection for long‑range identity preservation, and soft latent blending for edit locality. EditVid supports a wide range of editing tasks—including style transfer, attribute modification, object insertion, part‑level editing, and subject replacement—and outperforms the strongest training‑free baseline on FiVE while achieving competitive results on IVEBench, with a user study showing a 51.8% overall preference over seven competing methods.
By Adheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen, Muntasir Wahed, Nabeel Bashir, Xiaona Zhou, Tianjiao Yu, Vedant Shah, Ismini Lourentzou
Edit‑VAR is a training‑free, inversion‑free framework that uses a pretrained visual autoregressive video model for text‑guided video editing. It encodes the source video into multi‑scale discrete tokens and applies probability‑guided conditional token replacement, attention‑guided token‑wise and scale‑aware modulation, and scale‑decoupled generation to preserve source appearance while enabling precise edits. The method also includes residual‑guided token pruning to reduce inference cost, and experimental results show it outperforms existing training‑free video editing methods in fidelity, source preservation, temporal coherence, and efficiency.
By Chongbo Zhao, Jiangming Wang, Xilai Wang, Xinyu Wang, Jingyi Tang, Chunjie Hao, Pengjie Song, Yue Ma