OmniEdit-Bench introduces a comprehensive benchmark for instruction-based video editing (IVE), addressing limitations of existing datasets by covering spatial, temporal, audio, and reference-based editing tasks and distinguishing explicit from implicit instructions. The evaluation framework assesses editing quality across accuracy, preservation, realism, and consistency, using human judgments and vision-language models, and incorporates an accuracy-aware penalty to ensure instruction fidelity. Experiments reveal that current IVE models perform poorly, highlighting the need for improved methods.
By Chenxuan Miao, Yutong Feng, Yi Lu, Yunfeng Yan, Donglian Qi, Shiwei Zhang, Yu Liu, Xi Chen, Hengshuang Zhao
Timeline-Bench is a benchmark comprising 56 real video‑editing tasks that require AI agents to transform raw production material into finished videos. Each task includes a brief, source assets, a container, and a set of tests that assess format, content, brief compliance, and quality based on 2,582 blind judgments by 43 video editors. In evaluations, the best agent resolved only 15 of the 56 tasks, and most failures were due to quality tests rather than technical errors.
arXiv:2608. 08491v1 Announce Type: new Abstract: Reward models are a bottleneck for reinforcement learning in embodied AI.
By Yidong Wang, Yan Zhan, Ziteng Feng, Zhenyu Cui, Ziyi Zhou, Renzhao Liang, Jiaxuan Zhu, Zilei Yang, Yiran Zhao, Zhongkuan Mao, Bo Jia, Hanchu Ni, Chenggang Xie, Biao Liu, Yi Zhang, Yong Dai, Xiaozhu Ju, Wei Ye, Shikun Zhang
VideoGen-Agent is a multimodal agent that uses multitask agentic reinforcement learning to coordinate external tools for video generation. It learns to augment, generate, and verify videos through multi‑turn interactions, guided by prompts and intermediate observations. On the new VABench benchmark, the agent improves base text‑to‑video performance by 19.1 points, and further upgrades to generation tools raise the score to 86.1, with human raters favoring the upgraded configuration in 84.3% of comparisons.
By Binxu Li, Haoyi Duan, Yuhui Zhang, Yaohui Zhang, Zihao Lin, Kaituo Feng, Suozhi Huang, Xiangyi Li, Yu Li, Chunyuan Li, Shilong Liu, Mengdi Wang
arXiv:2607. 22632v1 Announce Type: new Abstract: The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans.
By Yexiang Liu, Wen Zhong, Sijie Zhu, Xin Gu, Fan Chen, Junxian Duan, Jie Cao, Longyin Wen, Zhenfang Chen
Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday users. However, editing models trained primarily for single-turn edits often break down in multi-turn editing--the natural interactive setting where a user iteratively refines an image based on the model's own previous outputs.
arXiv:2606. 00931v1 Announce Type: cross Abstract: Instruction-guided image editing is becoming a general interface for visual work, yet existing benchmarks still focus largely on narrow appearance edits and do not fully capture the diversity of real-image tasks in professional workflows.
By Fangzhou Lin, Peiran Li, Lingyu Xu, Wenjing Chen, Qianwen Ge, Shuo Xing, Mingyang Wu, Xiangbo Gao, Siyuan Yang, Kazunori Yamada, Ziming Zhang, Haichong Zhang, Zhen Dong, Ming-Hsuan Yang, Zhengzhong Tu
Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, sp...
KwaiMind is a commercial image editing system that combines general editing capabilities with e-commerce specialization. It uses an agent-based data engine with 1.8 million editing pairs and a multimodal diffusion transformer trained through pre‑training, fine‑tuning, preference optimization, and online reinforcement learning. The system is guided by a vision‑language judge and specialized rewards for click‑through rate, text rendering, and product consistency, and it achieves top scores on ImgEdit, GEdit, REDEdit, and the new Ecom‑Bench, while improving predicted and actual CTR in offline and online experiments.
By Junlong Wu, Zijun Li, Yuting Hu, Jia Sun, Pengcheng Wei, Yimin Zhou, Honglie Wang, Huaiqing Wang, Dewen Fan, Fei Zuo, Haixuan Gao, Lihui Peng, Tingxuan She, Yuqing Li, Boheng Zhang, Fan Yang, Wenwu Ou
arXiv:2608. 07565v1 Announce Type: cross Abstract: Conversational assistants increasingly recommend follow-up edits to help users continue a task.
By Zhijing Zhang, Jinpeng Yu, Xin Song, Bingnan Li, Chuyue Li, Changhui Du, Xiaolin Fang, Jiaming Liu, Ruihua Huang
arXiv:2606.08091v2 Announce Type: replace
Abstract: Agentic long video generation requires planning, tool orchestration, and cross-clip coordination over a long horizon. Most existing video agents ei...
By Jianhui Wei, Yan Zhang, Jie Tan, Hengchuan Zhu, Xiaotian Zhang, Ziyi Chen, Daoan Zhang, Wei Xu, Yeying Jin, Zuozhu Liu
AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation models~(AIVGMs) have become increasingly difficult to discern using conventional evaluation criteria, such as visual fidelity and semantic instruction following.