Vidu S1: A Real-Time Interactive Video Generation Model
arXiv:2607. 03118v1 Announce Type: cross Abstract: We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters.
Vidu S2 is a system that includes Vidu S2-Avatar, a real‑time interactive digital‑character model, and Vidu S2-Editing, a real‑time video editing model. It enables real‑time 720p video generation with dynamic references and improved instruction following, such as dancing, and allows real‑time editing of video streams for style rendering, clothing replacement, character replacement, and background replacement. Experiments show Vidu S2 outperforms all baselines, and a playable online demo is available at https://vidu.com/vidu-stream.
arXiv:2607. 03118v1 Announce Type: cross Abstract: We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters.
EditaLive! is a new real‑time framework for character video editing in live streaming, built on a pretrained image animation model (Wan‑Animate) that separates appearance from motion. It uses the CharEdit‑50K dataset for reference‑frame editing and video reconstruction, and adapts the model from offline bidirectional to causal streaming generation. A self‑rollout distillation strategy compresses the model into a two‑step sampler, employing fixed RoPE, alignment forcing, and first‑frame preserved sparse attention to reduce appearance drift and achieve low‑latency inference while preserving facial expressions.
EditStream is a unified DiT‑based framework that supports a wide range of interactive video tasks—Text‑to‑Video, Image‑to‑Video, Video‑to‑Video, Editing Propagation, Reference‑guided Video Editing, and Camera Pose Change—within a single system. It achieves fast, few‑step autoregressive generation by applying a two‑stage distillation process that combines Velocity Moment Matching with autoregressive unrolling, thereby preserving motion quality and temporal stability. The approach aims to make high‑quality diffusion‑based video models practical for real‑time creative workflows.
arXiv:2607.26694v3 Announce Type: replace Abstract: We present Visko Orbis 1.0, a Live Model for real-time, interactive long video generation. Users can change the prompt at any moment during generat...
arXiv:2609.40356v1 Announce Type: cross Abstract: Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits th...
Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration.
arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.
arXiv:2609.36598v1 Announce Type: new Abstract: A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is un...
arXiv:2609.13830v1 Announce Type: new Abstract: We present DiVA, a deeply interactive digital life simulator pioneering a new paradigm for long-term, open-ended interactive experiences within digital...
Memory-V2V is a memory‑augmented video‑to‑video diffusion framework designed to improve cross‑turn consistency in multi‑turn video editing. It stores previous outputs in an external memory, retrieves relevant edits, and incorporates them via relevance‑aware tokenization and adaptive compression, allowing scalable conditioning without linear computational growth. Experiments on iterative video novel view synthesis and text‑guided long video editing show that Memory‑V2V enhances consistency while preserving visual quality and outperforming strong baselines with modest overhead.
arXiv:2506. 10915v2 Announce Type: replace-cross Abstract: Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators.
TokenDial introduces a Visual Dial Space (V+) where the channel dimension of visual patch tokens in video diffusion transformers acts as a semantic control space. By learning additive directions in V+, the framework enables continuous slider-style edits for appearance and motion attributes without altering the pretrained generator. The method demonstrates improved controllability and content preservation compared to prior video editing techniques.