ID-V2V: Identity-Preserving Video Restylization
arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.
arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.
EditStream is a unified DiT‑based framework that supports a wide range of interactive video tasks—Text‑to‑Video, Image‑to‑Video, Video‑to‑Video, Editing Propagation, Reference‑guided Video Editing, and Camera Pose Change—within a single system. It achieves fast, few‑step autoregressive generation by applying a two‑stage distillation process that combines Velocity Moment Matching with autoregressive unrolling, thereby preserving motion quality and temporal stability. The approach aims to make high‑quality diffusion‑based video models practical for real‑time creative workflows.
Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration.
arXiv:2606. 19676v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success in image and video generation and editing.
arXiv:2607. 19895v1 Announce Type: cross Abstract: Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion.
arXiv:2607.18227v2 Announce Type: replace Abstract: In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and imag...
Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet the complex creative demands of real-world scenarios. To bridge this gap, we present Goku, a large-scale dataset featuring 2 million high-quality, instruction-aligned video editing pairs, which is the first to extend task boundaries from basic appearance editing to multi-task and structural manipulations(e.
arXiv:2601.16296v3 Announce Type: replace-cross Abstract: Video-to-video diffusion models achieve impressive single-turn editing performance, but practical editing workflows are inherently iterative....
arXiv:2606. 11751v1 Announce Type: cross Abstract: Multi-turn image editing is essential for iterative design, yet current models often struggle with identity drift and error accumulation over successive steps.
Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational overh...
LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. It preserves bounded bidirectional modeling within a fixed‑size window, emits clean video chunks iteratively, and uses two memory modules—a bounded temporal memory and a persistent global appearance memory—to sustain long‑term consistency. A progressive distillation process further aligns teacher‑based bidirectional learning with causal few‑step inference, resulting in superior generation quality with 26× lower latency and 11× higher throughput compared to comparable models.
RefVideo-6M is a new large-scale reference-guided editing dataset that includes 5 million video editing samples and 1 million image editing samples, each paired with about 6 million visual references. The dataset is constructed to avoid artifacts by using real, artifact‑free videos as targets and filtering input conditions with multiple editing experts, thereby providing reliable supervision. It enables models to learn fine‑grained visual correspondence beyond text‑only instructions and supports the training of a reference‑guided video editing model, Ref‑MoT, which shows improved visual quality, controllability, and reference consistency.