Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level or text-only video generation pipelines struggle to meet. We study plot-to-short-drama generation, where a global plot and local context are transformed into visually grounded multi-shot videos.
arXiv:2603.11421v2 Announce Type: replace
Abstract: Text-driven video generation has democratized film creation, but camera control in cinematic multi-shot scenarios remains a significant block. Impl...
By Songlin Yang, Zhe Wang, Xuyi Yang, Songchun Zhang, Xianghao Kong, Taiyi Wu, Xiaotong Zhao, Ran Zhang, Alan Zhao, Anyi Rao
CamPilot is a multi‑agent cinematic assistant that combines cinematographic planning with camera‑work control to generate more coherent and aesthetically pleasing movies from text prompts. It learns camera‑work planning from 14,000 professional films using a GRPO‑based learning paradigm, capturing motion patterns, composition principles, and cross‑shot relationships. The system is evaluated with a new benchmark, CamEval, and outperforms existing text‑to‑movie methods in cinematographic control and quality.
By Yang Wu, Stefano Petrangeli, Ishita Dasgupta, Yu Shen
arXiv:2609.38683v1 Announce Type: cross
Abstract: Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves...
By Ziqi Zhou, Yujian Yuan, Laura Sevilla-Lara
arXiv:2606. 26964v1 Announce Type: new Abstract: As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe.
By Jiaming Bian, Bingliang Li, Yuehao Wu, Pichao Wang, Zhi Wang, Hailan Ma, Huadong Mo, Zhenhong Sun
MVAgent is a multi‑agent pipeline for multi‑shot video generation that ensures consistent character appearance, stable spatial layout, and continuous character state across shots. The system uses typed conditioning inputs: a Spatial Grounding agent samples camera views, an Observer records shot endings into a continuity memory, a Transition agent builds action and spatial references for subsequent shots, and an Orchestrator composes these inputs into generator requests. Trained with agentic reinforcement learning (Trunk‑GDPO) while keeping the generator and judges frozen, MVAgent achieves the highest cross‑shot consistency and narrative‑planning quality on ViMax‑Bench and is preferred over the strongest agentic baseline in human evaluation.
By Xiangyu Kong, Wenjie Zhou, Fengping Tian, Lihua Fang, Haoqin Sun, Chenyang Lyu, Longyue Wang, Weihua Luo