FRAMEWORKERS is a task‑centric, multi‑agent framework designed for end‑to‑end AI‑generated video production. It uses a central Director to dynamically manage a task stack and an Assistant to execute tasks within a shared Workspace, leveraging modular sub‑agents that can be added without redesigning the workflow. The system is fine‑tuned with supervised learning and policy optimization, outperforming existing LLM planners and fixed pipelines in routing accuracy, failure recovery, and overall video quality.
By Zhendong Li, Lei Sun, Letian Shi, Deheng Zhang, Ruibo Ming, Mengshun Hu, Dannong Xu, Jian Wang, Danda Paudel, Luc Van Gool, Jinjin Gu
arXiv:2607. 01709v1 Announce Type: new Abstract: Agents are increasingly used to construct workflows and assist humans in completing recurring tasks more efficiently.
By Zongxia Li, Dawei Liu, Fuxiao Liu, Yuhang Zhou, Xiyang Wu, Jingxi Chen, Jing Xie, Xiaomin Wu, Lichao Sun
arXiv:2602. 11790v2 Announce Type: replace Abstract: Although recent end-to-end video generation models demonstrate impressive performance in visually oriented content creation, they remain limited in scenarios that require strict logical rigor and precise knowledge representation, such as instructional and educational media.
By Lingyong Yan, Jiulong Wu, Dong Xie, Weixian Shi, Deguo Xia, Jizhou Huang
VideoResearcher is a training‑free, multi‑agent framework that autonomously designs, tests, and refines high‑impact tools for long‑video understanding. It operates through dual Solving and Evolving loops, analyzing tool‑use trajectories to identify gaps, coordinating specialized agents to develop and validate executable tools, and reusing evolved tools to improve evidence acquisition in subsequent reasoning. The approach achieves state‑of‑the‑art performance among self‑improving agents and approaches the human‑designed upper bound, demonstrating a paradigm that expands agent capabilities while reducing costly manual engineering.
By Dingqiang Ye, Dongdi Zhao, Kaishen Wang, Qingqiao Hu, Jingchen Sun, Yijun Liang, Yuqi Jia, Yiqiao Huang, Yunjie Tian, Jiaxing Zhang, Chuanyang Jin, Ke Zhang, Vishal M. Patel, Di Fu
Timeline-Bench is a benchmark comprising 56 real video‑editing tasks that require AI agents to transform raw production material into finished videos. Each task includes a brief, source assets, a container, and a set of tests that assess format, content, brief compliance, and quality based on 2,582 blind judgments by 43 video editors. In evaluations, the best agent resolved only 15 of the 56 tasks, and most failures were due to quality tests rather than technical errors.
arXiv:2609.37950v1 Announce Type: new
Abstract: Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However,...
By Bingjun Luo, Jialin Guo, Siqi Li
Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, sp...
Long-form video understanding requires locating sparse, question-relevant evidence in long, multimodal videos. Real-world video distributions differ in modality-specific information density, content structure, and evidence patterns, causing fixed video-agent designs to incur redundant processing or fail when mismatched.
VideoGen-Agent is a multimodal agent that uses multitask agentic reinforcement learning to coordinate external tools for video generation. It learns to augment, generate, and verify videos through multi‑turn interactions, guided by prompts and intermediate observations. On the new VABench benchmark, the agent improves base text‑to‑video performance by 19.1 points, and further upgrades to generation tools raise the score to 86.1, with human raters favoring the upgraded configuration in 84.3% of comparisons.
By Binxu Li, Haoyi Duan, Yuhui Zhang, Yaohui Zhang, Zihao Lin, Kaituo Feng, Suozhi Huang, Xiangyi Li, Yu Li, Chunyuan Li, Shilong Liu, Mengdi Wang
arXiv:2609.40048v1 Announce Type: new
Abstract: Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget,...
By Yiduo Jia, Muzhi Zhu, Jinchuan Shi, Hao Zhong, Yuling Xi, Ke Liu, Hao Chen
VideoEvolve is a framework that automatically evolves agent harnesses for video temporal grounding, a task that localizes events in videos based on natural-language queries. It introduces a Cloze-Structured Harness Representation to keep stage interfaces stable while allowing agent workflows and instructions to change, and uses Branch-Guided Harness Evolution to preserve promising code branches, guide local edits with execution feedback, and validate improvements. Experiments show that this automated evolution improves grounding performance across multiple benchmarks, with instruction refinement consistently yielding gains.
By Bingjun Luo, Yuhuan Fan, Jialin Guo, Siqi Li
arXiv:2608. 14015v1 Announce Type: cross Abstract: Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time.
By Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang, Liang Chen, Jiahao Xie, Jiangtang Wang, Shuai Shao, Anchao Yang, Yutong Bai, Yan Wang