Veo 3.1 Ingredients to Video: More consistency, creativity and control
Our latest Veo update generates lively, dynamic clips that feel natural and engaging — and supports vertical video generation.
Introducing Veo 3 and Imagen 4, and a new tool for filmmaking called Flow.
Our latest Veo update generates lively, dynamic clips that feel natural and engaging — and supports vertical video generation.
SkillPE is a prompt‑engineering framework that evolves reusable cinematic skills from expert‑authored seeds to improve text‑to‑video generation for non‑experts. It encodes shot logic, composition, lighting, sound design, and other filmmaking cues in a fine‑grained format, and uses movie references classified as resonators, dissonants, and divergents to refine skill application and inspire creative alternatives. Experiments on StoryEval and VBench demonstrate up to 1.40‑point gains over the strongest baseline and 0.51 points over seed skills on a 7‑point four‑dimensional evaluation, while remaining competitive on benchmark‑native metrics.
Despite remarkable progress in text-guided image editing, generative models frequently fail to preserve visual object consistency, defined as the preservation of a subject's key attributes throughout the editing process. We address this limitation through three contributions.
The paper "Iterative Flow Matching: Path Correction and Gradual Refinement for Enhanced Generative Modeling" investigates the use of flow matching for image generation and identifies that this approach can produce hallucinations—unrealistic images. It proposes an iterative refinement process that can be incorporated into virtually any generative modeling technique to improve performance and robustness. The authors demonstrate how their method corrects the generation path and gradually refines outputs to mitigate hallucinations.
We partnered with Darren Aronofsky, Eliza McNitt and a team of more than 200 people to make a film using Veo and live-action filmmaking.
arXiv:2608. 08101v1 Announce Type: new Abstract: Generative AI has emerged as one of the most transformative forces in modern artificial intelligence, reshaping how we create, imagine, and interact with digital content.
We’re rolling out significant updates to Veo that give people even more creative control.
CamPilot is a multi‑agent cinematic assistant that combines cinematographic planning with camera‑work control to generate more coherent and aesthetically pleasing movies from text prompts. It learns camera‑work planning from 14,000 professional films using a GRPO‑based learning paradigm, capturing motion patterns, composition principles, and cross‑shot relationships. The system is evaluated with a new benchmark, CamEval, and outperforms existing text‑to‑movie methods in cinematographic control and quality.
arXiv:2607.
The paper compares diffusion and rectified flow objectives within the MotionGPT3 framework for text-driven motion generation. Experiments on HumanML3D show that rectified flow converges faster, achieves strong test performance earlier, and matches or exceeds diffusion quality while requiring fewer sampling steps. The study isolates the generative objective’s impact, demonstrating that rectified flow’s benefits transfer to continuous-latent motion generation.
The paper presents a method for generating 3D staging—human poses, lighting, and camera setup—directly from affective textual descriptions. It builds a dataset of 11,911 text–staging pairs derived from 2,328 figurative paintings, reconstructing SMPL bodies, estimating illumination, and recovering camera parameters. A flow‑matching transformer is trained to produce variable‑size scenes and multiple staging alternatives, achieving a 32.2% retrieval R@1 on held‑out prompts, outperforming a CLIP‑based baseline.
arXiv:2606. 27377v1 Announce Type: cross Abstract: Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing.