FRAMEWORKERS is a task‑centric, multi‑agent framework designed for end‑to‑end AI‑generated video production. It uses a central Director to dynamically manage a task stack and an Assistant to execute tasks within a shared Workspace, leveraging modular sub‑agents that can be added without redesigning the workflow. The system is fine‑tuned with supervised learning and policy optimization, outperforming existing LLM planners and fixed pipelines in routing accuracy, failure recovery, and overall video quality.
By Zhendong Li, Lei Sun, Letian Shi, Deheng Zhang, Ruibo Ming, Mengshun Hu, Dannong Xu, Jian Wang, Danda Paudel, Luc Van Gool, Jinjin Gu
WeAgent-MMGenEdit is a comprehensive framework for multimodal agentic image generation and editing that addresses the unreliability of current models when prompts require external world knowledge. It introduces a multimodal harness with persistent evidence management, a scalable data construction pipeline producing 23K supervised trajectories and 14.7K RL tasks, and a bilingual benchmark (WeBench-MMGenEdit) for knowledge-intensive generation and multi-image editing. Post‑training methods based on SFT and RL further refine the agent policy and image backend, enabling a 30B‑parameter policy to outperform similarly sized models and approach the performance of a 1T‑parameter agent.
By Hui Zhang, Zongkai Liu, Liqiang Niu, Juntao Liu, Han Li, Zhen Cao, Wenchao Chen, Chengduo Zhao, Fandong Meng
The paper introduces the Multi-Instruction Multi-Shot Long-Video Editing (MMLVE) task, aiming to enable consistent editing of long videos with multiple instructions. It proposes an agentic editing framework that combines Large Language Models and Vision-Language Models for shot-level decoupling and precise instruction parsing. The authors also present the MMLVE-Bench dataset and evaluation metrics, showing that their MMLVE-Agent outperforms existing state‑of‑the‑art methods by eliminating hallucinations and preserving temporal consistency.
By Chenyang Wu, Fuchen Long, Binyuan Huang, Xinlong Sun, Xi Chen, Chun-Le Guo, Chongyi Li
arXiv:2608. 12290v1 Announce Type: cross Abstract: Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows.
By Aman Tyagi, Hemanth Boinpally, Jonathan Chen, Douglas Gebert, Steven Hickson
The paper introduces the Multi-Instruction Multi-Shot Long-Video Editing (MMLVE) task, aiming to edit long videos with multiple instructions while maintaining consistency across shots, decoupling instructions, and preserving spatiotemporal structure. It proposes an agentic editing framework that combines Large Language Models and Vision‑Language Models for shot-level decoupling and precise instruction parsing. A new dataset, MMLVE‑Bench, and three evaluation metrics are created to benchmark this task, and experiments show the proposed MMLVE‑Agent outperforms existing closed‑source state‑of‑the‑art methods by eliminating hallucinations and ensuring seamless transitions.
Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, sp...