DeepMind Blog

Fuel your creativity with new generative media models and tools

Introducing Veo 3 and Imagen 4, and a new tool for filmmaking called Flow.

Hugging Face Trending Papers
6d ago

SkillPE: Creativity-Oriented Cinematic Skill Evolution for Text-to-Video Prompt Engineering

SkillPE is a prompt‑engineering framework that evolves reusable cinematic skills from expert‑authored seeds to improve text‑to‑video generation for non‑experts. It encodes shot logic, composition, lighting, sound design, and other filmmaking cues in a fine‑grained format, and uses movie references classified as resonators, dissonants, and divergents to refine skill application and inspire creative alternatives. Experiments on StoryEval and VBench demonstrate up to 1.40‑point gains over the strongest baseline and 0.51 points over seed skills on a 7‑point four‑dimensional evaluation, while remaining competitive on benchmark‑native metrics.

arXiv AI
Aug 20

Iterative Flow Matching: Path Correction and Gradual Refinement for Enhanced Generative Modeling

The paper "Iterative Flow Matching: Path Correction and Gradual Refinement for Enhanced Generative Modeling" investigates the use of flow matching for image generation and identifies that this approach can produce hallucinations—unrealistic images. It proposes an iterative refinement process that can be incorporated into virtually any generative modeling technique to improve performance and robustness. The authors demonstrate how their method corrects the generation path and gradually refines outputs to mitigate hallucinations.

By Eldad Haber, Shadab Ahamed, Md. Shahriar Rahim Siddiqui, Niloufar Zakariaei, Moshe Eliasof
arXiv Computer Vision
Sep 11

CamPilot: A Multi-Agent Cinematic Assistant for Camera-Controlled Movie Generation

CamPilot is a multi‑agent cinematic assistant that combines cinematographic planning with camera‑work control to generate more coherent and aesthetically pleasing movies from text prompts. It learns camera‑work planning from 14,000 professional films using a GRPO‑based learning paradigm, capturing motion patterns, composition principles, and cross‑shot relationships. The system is evaluated with a new benchmark, CamEval, and outperforms existing text‑to‑movie methods in cinematographic control and quality.

By Yang Wu, Stefano Petrangeli, Ishita Dasgupta, Yu Shen
arXiv Machine Learning
Aug 19

From Diffusion to Flow: Efficient Motion Generation in MotionGPT3

The paper compares diffusion and rectified flow objectives within the MotionGPT3 framework for text-driven motion generation. Experiments on HumanML3D show that rectified flow converges faster, achieves strong test performance earlier, and matches or exceeds diffusion quality while requiring fewer sampling steps. The study isolates the generative objective’s impact, demonstrating that rectified flow’s benefits transfer to continuous-latent motion generation.

By Jaymin Bhan, JiHong Jeon, SangYeop Jeong
arXiv AI
Sep 16

Text-Driven Artistic Staging: 3D Posing, Lighting, and Camera References from Paintings

The paper presents a method for generating 3D staging—human poses, lighting, and camera setup—directly from affective textual descriptions. It builds a dataset of 11,911 text–staging pairs derived from 2,328 figurative paintings, reconstructing SMPL bodies, estimating illumination, and recovering camera parameters. A flow‑matching transformer is trained to produce variable‑size scenes and multiple staging alternatives, achieving a 32.2% retrieval R@1 on held‑out prompts, outperforming a CLIP‑based baseline.

By Yunge Wen
arXiv Machine Learning
Jun 26

DanceOPD: On-Policy Generative Field Distillation

arXiv:2606. 27377v1 Announce Type: cross Abstract: Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing.

By Wei Zhou, Xiongwei Zhu, Zelin Xu, Bo Dong, Lixue Gong, Yongyuan Liang, Meng Chu, Leigang Qu, Lingdong Kong, Wei Liu, Tat-Seng Chua