Veo 3.1 Ingredients to Video: More consistency, creativity and control
Our latest Veo update generates lively, dynamic clips that feel natural and engaging — and supports vertical video generation.
Introducing Veo 3 and Imagen 4, and a new tool for filmmaking called Flow.
Our latest Veo update generates lively, dynamic clips that feel natural and engaging — and supports vertical video generation.
Despite remarkable progress in text-guided image editing, generative models frequently fail to preserve visual object consistency, defined as the preservation of a subject's key attributes throughout the editing process. We address this limitation through three contributions.
We partnered with Darren Aronofsky, Eliza McNitt and a team of more than 200 people to make a film using Veo and live-action filmmaking.
arXiv:2608. 08101v1 Announce Type: new Abstract: Generative AI has emerged as one of the most transformative forces in modern artificial intelligence, reshaping how we create, imagine, and interact with digital content.
We’re rolling out significant updates to Veo that give people even more creative control.
arXiv:2607.
arXiv:2603. 26747v3 Announce Type: replace-cross Abstract: Recent text-driven motion generation methods span both discrete token-based approaches and continuous-latent formulations.
arXiv:2606. 27377v1 Announce Type: cross Abstract: Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing.
arXiv:2607. 24241v1 Announce Type: cross Abstract: Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models.
At OpenAI, we have long believed image generation should be a primary capability of our language models. That’s why we’ve built our most advanced image generator yet into GPT‑4o.
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft.
Prior work on aesthetic composition typically produces a single aesthetically pleasing crop, overlooking the narrative value of composing multiple shots from one scene. In practice, multi-shot composition is critical for downstream creative workflows: commercial posters often require multiple crops with different emphases (e.