Current video editors can insert objects but often struggle to make them participate in interactions such as being picked up or manipulated. We introduce ALIVE, a framework that makes inserted objects...
arXiv:2606. 08415v1 Announce Type: cross Abstract: While recent text-guided video editing models excel at elementary tasks (e.
By Jiangtao Wu, Jiaming Wang, Yiwen He, Yuanxing Zhang, Shihao Li, Dunyuan Liu, Xuedong Zhao, Jialu Chen, Zekun Moore Wang, Jiaheng Liu
OmniEdit-Bench introduces a comprehensive benchmark for instruction-based video editing (IVE), addressing limitations of existing datasets by covering spatial, temporal, audio, and reference-based editing tasks and distinguishing explicit from implicit instructions. The evaluation framework assesses editing quality across accuracy, preservation, realism, and consistency, using human judgments and vision-language models, and incorporates an accuracy-aware penalty to ensure instruction fidelity. Experiments reveal that current IVE models perform poorly, highlighting the need for improved methods.
By Chenxuan Miao, Yutong Feng, Yi Lu, Yunfeng Yan, Donglian Qi, Shiwei Zhang, Yu Liu, Xi Chen, Hengshuang Zhao
arXiv:2603. 06140v2 Announce Type: replace-cross Abstract: Video object insertion is fundamental to video editing, yet existing diffusion methods often produce visually plausible but physically inconsistent results.
By Bohai Gu, Taiyi Wu, Dazhao Du, Jian Liu, Shuai Yang, Xiaotong Zhao, Alan Zhao, Song Guo
Memory-V2V is a memory‑augmented video‑to‑video diffusion framework designed to improve cross‑turn consistency in multi‑turn video editing. It stores previous outputs in an external memory, retrieves relevant edits, and incorporates them via relevance‑aware tokenization and adaptive compression, allowing scalable conditioning without linear computational growth. Experiments on iterative video novel view synthesis and text‑guided long video editing show that Memory‑V2V enhances consistency while preserving visual quality and outperforming strong baselines with modest overhead.
By Dohun Lee, Chun-Hao Paul Huang, Xuelin Chen, Jong Chul Ye, Duygu Ceylan, Hyeonho Jeong
arXiv:2506.01004v3 Announce Type: replace-cross
Abstract: Unlike traditional video editing or inpainting, video semantic mixing fuses a reference concept with a moving target entity to produce a hybr...
By Tong Zhang, Victor Escorcia, Juan C Leon Alcazar, Bernard Ghanem
arXiv:2602.08277v3 Announce Type: replace-cross
Abstract: The landscape of AI video generation is undergoing a pivotal shift: moving beyond general generation - which relies on exhaustive prompt-engi...
By Xiangbo Gao, Renjie Li, Xinghao Chen, Yuheng Wu, Suofei Feng, Qing Yin, Zhengzhong Tu
RefVideo-6M is a new large-scale reference-guided editing dataset that includes 5 million video editing samples and 1 million image editing samples, each paired with about 6 million visual references. The dataset is constructed to avoid artifacts by using real, artifact‑free videos as targets and filtering input conditions with multiple editing experts, thereby providing reliable supervision. It enables models to learn fine‑grained visual correspondence beyond text‑only instructions and supports the training of a reference‑guided video editing model, Ref‑MoT, which shows improved visual quality, controllability, and reference consistency.
By Bojia Zi, Xiaoyan Yang, Yu Zhou, Ruijie Sun, Lihan Zhang, Bin Liang, Kam-Fai Wong, Haibin Huang, Chi Zhang, Xuelong Li
arXiv:2604. 14556v2 Announce Type: replace-cross Abstract: Video object insertion places a user-specified object in an existing dynamic scene.
By Qi Xia, Peishan Cong, Yichen Yao, Ziyi Wang, Yaoqin Ye, Yuexin Ma
The paper presents a framework for creating a scalable synthetic dataset of controllable video interactions by generating explicit start and end state images using image editing models. It introduces State‑Guided Sampling (SGS) to produce seamless videos anchored on these states, reducing artifacts seen in naive conditional generation. An automated evaluation system aligned with human judgments is also developed, and experiments demonstrate that fine‑tuning a base model on this dataset markedly improves its ability to generate plausible interactions.
By Jiho Jang, Jinyoung Kim, Nojun Kwak, Kyungjune Kim
arXiv:2607.18227v2 Announce Type: replace
Abstract: In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and imag...
By Dingyun Zhang, Lixue Gong, Wei Liu
CameraEditor is a new framework that transforms camera-controlled image editing into a temporal sequence prediction problem. By using video diffusion models, it incorporates a geometric perception module and dynamic reference routing to create precise visual references through dynamic panorama cropping. The method also inserts intermediate transition frames to handle large perspective shifts, maintaining content identity and spatial coherence, and is evaluated on a dataset of 5,760 instances with a benchmark of 462 test cases, achieving state‑of‑the‑art performance.
By Xin Shen, Chengyou Jia, Keshuo Xing, Zifeng Zhu, Changliang Xia, Bowen Ping, Zhuohang Dang, Hangwei Qian, Minnan Luo