arXiv AI By Hengji Zhou, Sijie Liu, Jianrun Chen, Xingchen Zou, Lianghao Xia, Liqiang Nie

DramaDirector: Geometry-Guided Short Drama Generation

Read the original on arXiv AI →

arXiv:2606. 24107v1 Announce Type: cross Abstract: Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level or text-only video generation pipelines struggle to meet.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 11

CamPilot: A Multi-Agent Cinematic Assistant for Camera-Controlled Movie Generation

CamPilot is a multi‑agent cinematic assistant that combines cinematographic planning with camera‑work control to generate more coherent and aesthetically pleasing movies from text prompts. It learns camera‑work planning from 14,000 professional films using a GRPO‑based learning paradigm, capturing motion patterns, composition principles, and cross‑shot relationships. The system is evaluated with a new benchmark, CamEval, and outperforms existing text‑to‑movie methods in cinematographic control and quality.

By Yang Wu, Stefano Petrangeli, Ishita Dasgupta, Yu Shen
arXiv Computer Vision
6d ago

MVAgent: Multi-Agent Video Generation via Consistent Condition Construction and Shot-Level Policy Optimization

MVAgent is a multi‑agent pipeline for multi‑shot video generation that ensures consistent character appearance, stable spatial layout, and continuous character state across shots. The system uses typed conditioning inputs: a Spatial Grounding agent samples camera views, an Observer records shot endings into a continuity memory, a Transition agent builds action and spatial references for subsequent shots, and an Orchestrator composes these inputs into generator requests. Trained with agentic reinforcement learning (Trunk‑GDPO) while keeping the generator and judges frozen, MVAgent achieves the highest cross‑shot consistency and narrative‑planning quality on ViMax‑Bench and is preferred over the strongest agentic baseline in human evaluation.

By Xiangyu Kong, Wenjie Zhou, Fengping Tian, Lihua Fang, Haoqin Sun, Chenyang Lyu, Longyue Wang, Weihua Luo