arXiv Computation and Language

AgenticGen: Reward-Guided Agentic Video Generation for Advertising

AgenticGen is a reward‑guided framework for generating advertising videos that splits the task into strategy selection and draft generation, allowing online business feedback to supervise each stage. It learns performance‑based rewards from accumulated online metrics and rubric‑based rewards aligned with human quality standards, then applies DPO and GRPO to refine policies. Offline tests confirm the reward models, and online A/B tests on TikTok show significant gains in CTR, CVR, and advertising value over a baseline.

arXiv Computer Vision
Sep 22

VideoGen-Agent: Reinforcing Video Generation Agents

VideoGen-Agent is a multimodal agent that uses multitask agentic reinforcement learning to coordinate external tools for video generation. It learns to augment, generate, and verify videos through multi‑turn interactions, guided by prompts and intermediate observations. On the new VABench benchmark, the agent improves base text‑to‑video performance by 19.1 points, and further upgrades to generation tools raise the score to 86.1, with human raters favoring the upgraded configuration in 84.3% of comparisons.

By Binxu Li, Haoyi Duan, Yuhui Zhang, Yaohui Zhang, Zihao Lin, Kaituo Feng, Suozhi Huang, Xiangyi Li, Yu Li, Chunyuan Li, Shilong Liu, Mengdi Wang
arXiv AI
Aug 25

Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation

arXiv:2608.21425v1 Announce Type: cross Abstract: Video generation is central to AI-powered content creation. Aligning generated videos with human preferences is a key criterion for evaluating genera...

By Nai-Xin Zhai, Weihua Cheng, Dexu Yu, Yikai Gu, Hanwen Du, Junchen Fu, Chenxi Huang, Yingwei Song, Liyuan Lillian Ma, Yang Ran, Youhua Li, Yongxin Ni
arXiv Computer Vision
4d ago

Think Before You Score: Thinking Reward Model for Visual Generation

arXiv:2609.37372v1 Announce Type: new Abstract: Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and can...

By Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
arXiv AI
Jun 2

From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models

arXiv:2606. 00083v1 Announce Type: cross Abstract: Reinforcement learning relies on accurate reward functions, which are often hand-crafted or even unavailable in real-world applications, such as robotics.

By Christian Gumbsch, Leonardo Barcellona, Lennard Sch\"unemann, Platon Karageorgis, Andrii Zadaianchuk, Zehao Wang, Sergey Zakharov, Fabien Despinoy, Rahaf Aljundi, Efstratios Gavves
arXiv AI
Jun 2

Video Reasoning without Training

arXiv:2510. 17045v2 Announce Type: replace-cross Abstract: Video reasoning using Large Multimodal Models (LMMs) relies on costly reinforcement learning (RL) and verbose chain-of-thought, resulting in substantial computational overhead during both training and inference.

By Deepak Sridhar, Kartikeya Bhardwaj, Jeya Pradha Jeyaraj, Nuno Vasconcelos, Ankita Nayak, Harris Teague