Streaming Video Editing with Easy Adaptation
arXiv:2609.24788v1 Announce Type: new Abstract: In this paper, we propose SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion model but supports high-qualit...
Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration.
arXiv:2609.24788v1 Announce Type: new Abstract: In this paper, we propose SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion model but supports high-qualit...
EditStream is a unified DiT‑based framework that supports a wide range of interactive video tasks—Text‑to‑Video, Image‑to‑Video, Video‑to‑Video, Editing Propagation, Reference‑guided Video Editing, and Camera Pose Change—within a single system. It achieves fast, few‑step autoregressive generation by applying a two‑stage distillation process that combines Velocity Moment Matching with autoregressive unrolling, thereby preserving motion quality and temporal stability. The approach aims to make high‑quality diffusion‑based video models practical for real‑time creative workflows.
arXiv:2607. 19895v1 Announce Type: cross Abstract: Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion.
EditaLive! is a new real‑time framework for character video editing in live streaming, built on a pretrained image animation model (Wan‑Animate) that separates appearance from motion. It uses the CharEdit‑50K dataset for reference‑frame editing and video reconstruction, and adapts the model from offline bidirectional to causal streaming generation. A self‑rollout distillation strategy compresses the model into a two‑step sampler, employing fixed RoPE, alignment forcing, and first‑frame preserved sparse attention to reduce appearance drift and achieve low‑latency inference while preserving facial expressions.
arXiv:2609.37925v1 Announce Type: cross Abstract: Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Trai...
Memory-V2V is a memory‑augmented video‑to‑video diffusion framework designed to improve cross‑turn consistency in multi‑turn video editing. It stores previous outputs in an external memory, retrieves relevant edits, and incorporates them via relevance‑aware tokenization and adaptive compression, allowing scalable conditioning without linear computational growth. Experiments on iterative video novel view synthesis and text‑guided long video editing show that Memory‑V2V enhances consistency while preserving visual quality and outperforming strong baselines with modest overhead.
arXiv:2601. 09881v2 Announce Type: replace-cross Abstract: Large video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interactive applications remains limited due to their inefficient multi-step sampling process.
ViRDM is a new post‑training method for few‑step causal video generation that eliminates the need for a large teacher model and an online critic. By applying representation distribution matching (RDM) with a precomputed target distribution, a lightweight VAE decoder, and staged vector–Jacobian products, ViRDM overcomes memory, optimization, and temporal dynamics challenges. The approach reduces GPU memory usage and training time, achieving state‑of‑the‑art VBench performance with only 20 generator updates and 16 A100 GPU‑hours.
arXiv:2607. 18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly.
Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods are restricted to forward temporal generation, whereas practical video creation often requires flexible generation order, e.
Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training methods predominantly rely on Distribution Matching Distillation (DMD), requiring both...
arXiv:2607. 24560v1 Announce Type: cross Abstract: We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D.