FRAMEWORKERS is a task‑centric, multi‑agent framework designed for end‑to‑end AI‑generated video production. It uses a central Director to dynamically manage a task stack and an Assistant to execute tasks within a shared Workspace, leveraging modular sub‑agents that can be added without redesigning the workflow. The system is fine‑tuned with supervised learning and policy optimization, outperforming existing LLM planners and fixed pipelines in routing accuracy, failure recovery, and overall video quality.
By Zhendong Li, Lei Sun, Letian Shi, Deheng Zhang, Ruibo Ming, Mengshun Hu, Dannong Xu, Jian Wang, Danda Paudel, Luc Van Gool, Jinjin Gu
WeAgent-MMGenEdit is a comprehensive framework for multimodal agentic image generation and editing that addresses the unreliability of current models when prompts require external world knowledge. It introduces a multimodal harness with persistent evidence management, a scalable data construction pipeline producing 23K supervised trajectories and 14.7K RL tasks, and a bilingual benchmark (WeBench-MMGenEdit) for knowledge-intensive generation and multi-image editing. Post‑training methods based on SFT and RL further refine the agent policy and image backend, enabling a 30B‑parameter policy to outperform similarly sized models and approach the performance of a 1T‑parameter agent.
By Hui Zhang, Zongkai Liu, Liqiang Niu, Juntao Liu, Han Li, Zhen Cao, Wenchao Chen, Chengduo Zhao, Fandong Meng
The paper introduces the Multi-Instruction Multi-Shot Long-Video Editing (MMLVE) task, aiming to enable consistent editing of long videos with multiple instructions. It proposes an agentic editing framework that combines Large Language Models and Vision-Language Models for shot-level decoupling and precise instruction parsing. The authors also present the MMLVE-Bench dataset and evaluation metrics, showing that their MMLVE-Agent outperforms existing state‑of‑the‑art methods by eliminating hallucinations and preserving temporal consistency.
By Chenyang Wu, Fuchen Long, Binyuan Huang, Xinlong Sun, Xi Chen, Chun-Le Guo, Chongyi Li
arXiv:2608. 12290v1 Announce Type: cross Abstract: Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows.
By Aman Tyagi, Hemanth Boinpally, Jonathan Chen, Douglas Gebert, Steven Hickson
The paper introduces the Multi-Instruction Multi-Shot Long-Video Editing (MMLVE) task, aiming to edit long videos with multiple instructions while maintaining consistency across shots, decoupling instructions, and preserving spatiotemporal structure. It proposes an agentic editing framework that combines Large Language Models and Vision‑Language Models for shot-level decoupling and precise instruction parsing. A new dataset, MMLVE‑Bench, and three evaluation metrics are created to benchmark this task, and experiments show the proposed MMLVE‑Agent outperforms existing closed‑source state‑of‑the‑art methods by eliminating hallucinations and ensuring seamless transitions.
Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, sp...
VideoGen-Agent is a multimodal agent that uses multitask agentic reinforcement learning to coordinate external tools for video generation. It learns to augment, generate, and verify videos through multi‑turn interactions, guided by prompts and intermediate observations. On the new VABench benchmark, the agent improves base text‑to‑video performance by 19.1 points, and further upgrades to generation tools raise the score to 86.1, with human raters favoring the upgraded configuration in 84.3% of comparisons.
By Binxu Li, Haoyi Duan, Yuhui Zhang, Yaohui Zhang, Zihao Lin, Kaituo Feng, Suozhi Huang, Xiangyi Li, Yu Li, Chunyuan Li, Shilong Liu, Mengdi Wang
Bernini proposes a unified framework that separates semantic planning and pixel rendering for video generation and editing. An MLLM-based planner predicts target semantics in ViT embedding space, while a DiT-based renderer synthesizes pixels conditioned on this plan, text features, and source VAE features for editing. The approach introduces Segment-Aware 3D Rotary Positional Embedding and chain-of-thought reasoning, achieving state‑of‑the‑art performance on diverse video benchmarks.
By Bernini Team, Chenchen Liu, Junyi Chen, Lei Li, Lu Chi, Mingzhen Sun, Zhuoying Li, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge Bai, Zehuan Yuan
arXiv:2609.08275v1 Announce Type: new
Abstract: Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherence does not imply the ability to execute e...
By Tianyi Zeng, Junchao Liao, Yujie Wei, Ziying Zhang, Litao Li, Tianyi Wang, Zhichao Wei, Shuyao Xu, Wenwen Qiang, Siyu Zhu, Zhenghao Zhang, Long Qin
arXiv:2606.08091v2 Announce Type: replace
Abstract: Agentic long video generation requires planning, tool orchestration, and cross-clip coordination over a long horizon. Most existing video agents ei...
By Jianhui Wei, Yan Zhang, Jie Tan, Hengchuan Zhu, Xiaotian Zhang, Ziyi Chen, Daoan Zhang, Wei Xu, Yeying Jin, Zuozhu Liu
Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity causes minor variations in textual prompts or hyperparameters to yield drastically different outputs often necessitating inefficient, brute-force trial-and-error processes.
arXiv:2602. 11790v2 Announce Type: replace Abstract: Although recent end-to-end video generation models demonstrate impressive performance in visually oriented content creation, they remain limited in scenarios that require strict logical rigor and precise knowledge representation, such as instructional and educational media.
By Lingyong Yan, Jiulong Wu, Dong Xie, Weixian Shi, Deguo Xia, Jizhou Huang