ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
ALIVE is a new framework for first‑frame‑guided video editing that enables inserted objects to participate in coherent interactions with the source video, such as being picked up or manipulated. The authors curated 35,800 editing pairs from 3D‑rendered, model‑generated, and real‑world videos, and trained a vision‑language model to predict interaction guidance from the edited first frame and an object name. ALIVE outperforms the strongest baseline by 43.9% on overall performance and 4.4% on a general video object insertion benchmark, with VLM‑predicted guidance further improving interaction fidelity.
arXiv:2606. 08415v1 Announce Type: cross Abstract: While recent text-guided video editing models excel at elementary tasks (e.
OmniEdit-Bench introduces a comprehensive benchmark for instruction-based video editing (IVE), addressing limitations of existing datasets by covering spatial, temporal, audio, and reference-based editing tasks and distinguishing explicit from implicit instructions. The evaluation framework assesses editing quality across accuracy, preservation, realism, and consistency, using human judgments and vision-language models, and incorporates an accuracy-aware penalty to ensure instruction fidelity. Experiments reveal that current IVE models perform poorly, highlighting the need for improved methods.
Memory-V2V is a memory‑augmented video‑to‑video diffusion framework designed to improve cross‑turn consistency in multi‑turn video editing. It stores previous outputs in an external memory, retrieves relevant edits, and incorporates them via relevance‑aware tokenization and adaptive compression, allowing scalable conditioning without linear computational growth. Experiments on iterative video novel view synthesis and text‑guided long video editing show that Memory‑V2V enhances consistency while preserving visual quality and outperforming strong baselines with modest overhead.
arXiv:2603. 06140v2 Announce Type: replace-cross Abstract: Video object insertion is fundamental to video editing, yet existing diffusion methods often produce visually plausible but physically inconsistent results.
arXiv:2506.01004v3 Announce Type: replace-cross Abstract: Unlike traditional video editing or inpainting, video semantic mixing fuses a reference concept with a moving target entity to produce a hybr...