Introducing Veo 3.1 and advanced creative capabilities
We’re rolling out significant updates to Veo that give people even more creative control.
Our latest Veo update generates lively, dynamic clips that feel natural and engaging — and supports vertical video generation.
We’re rolling out significant updates to Veo that give people even more creative control.
Introducing Veo 3 and Imagen 4, and a new tool for filmmaking called Flow.
Vidu S2 is a system that includes Vidu S2-Avatar, a real‑time interactive digital‑character model, and Vidu S2-Editing, a real‑time video editing model. It enables real‑time 720p video generation with dynamic references and improved instruction following, such as dancing, and allows real‑time editing of video streams for style rendering, clothing replacement, character replacement, and background replacement. Experiments show Vidu S2 outperforms all baselines, and a playable online demo is available at https://vidu.com/vidu-stream.
arXiv:2607.26694v3 Announce Type: replace Abstract: We present Visko Orbis 1.0, a Live Model for real-time, interactive long video generation. Users can change the prompt at any moment during generat...
TokenDial introduces a Visual Dial Space (V+) where the channel dimension of visual patch tokens in video diffusion transformers acts as a semantic control space. By learning additive directions in V+, the framework enables continuous slider-style edits for appearance and motion attributes without altering the pretrained generator. The method demonstrates improved controllability and content preservation compared to prior video editing techniques.
arXiv:2511.20186v2 Announce Type: replace Abstract: Foundation video generation models such as WAN 2.2 exhibit strong text- and image-conditioned synthesis abilities but remain constrained to the sam...
arXiv:2608.23090v1 Announce Type: new Abstract: Looping videos are essential for practical applications such as web graphics, game development, and social media. However, existing approaches typicall...
WanPE is a 397‑B parameter prompt‑enhancement model that learns director‑level cinematic planning from 1.05 M real‑world videos. It generates shot‑level cinematic plans through video‑grounded reverse construction and uses Semantic‑Consistency GRPO (SC‑GRPO) to maintain user intent across shots and time. In evaluations, WanPE improves human preference over raw prompts by up to 50.86 points for 30‑second videos and outperforms commercial offerings for shorter durations.
We partnered with Darren Aronofsky, Eliza McNitt and a team of more than 200 people to make a film using Veo and live-action filmmaking.
Precise 3D spatial orchestration in text-to-video generation remains a significant challenge, particularly for multi-object scenes where semantic layout and temporal dynamics are often entangled. While existing depth-conditioned models achieve good structural fidelity, they necessitate dense, frame-accurate guidance that is labor-intensive to author for dynamic events involving deformable objects.
arXiv:2601. 08828v2 Announce Type: replace-cross Abstract: Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood.