arXiv Computer Vision By Zhenghong Zhou, Zhe Lin, Jiebo Luo, Yuqian Zhou

ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing

Read the original on arXiv Computer Vision →

ALIVE is a new framework for first‑frame‑guided video editing that enables inserted objects to participate in coherent interactions with the source video, such as being picked up or manipulated. The authors curated 35,800 editing pairs from 3D‑rendered, model‑generated, and real‑world videos, and trained a vision‑language model to predict interaction guidance from the edited first frame and an object name. ALIVE outperforms the strongest baseline by 43.9% on overall performance and 4.4% on a general video object insertion benchmark, with VLM‑predicted guidance further improving interaction fidelity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 3

OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing

OmniEdit-Bench introduces a comprehensive benchmark for instruction-based video editing (IVE), addressing limitations of existing datasets by covering spatial, temporal, audio, and reference-based editing tasks and distinguishing explicit from implicit instructions. The evaluation framework assesses editing quality across accuracy, preservation, realism, and consistency, using human judgments and vision-language models, and incorporates an accuracy-aware penalty to ensure instruction fidelity. Experiments reveal that current IVE models perform poorly, highlighting the need for improved methods.

By Chenxuan Miao, Yutong Feng, Yi Lu, Yunfeng Yan, Donglian Qi, Shiwei Zhang, Yu Liu, Xi Chen, Hengshuang Zhao
arXiv Machine Learning
Aug 27

Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing

Memory-V2V is a memory‑augmented video‑to‑video diffusion framework designed to improve cross‑turn consistency in multi‑turn video editing. It stores previous outputs in an external memory, retrieves relevant edits, and incorporates them via relevance‑aware tokenization and adaptive compression, allowing scalable conditioning without linear computational growth. Experiments on iterative video novel view synthesis and text‑guided long video editing show that Memory‑V2V enhances consistency while preserving visual quality and outperforming strong baselines with modest overhead.

By Dohun Lee, Chun-Hao Paul Huang, Xuelin Chen, Jong Chul Ye, Duygu Ceylan, Hyeonho Jeong