arXiv AI By Zhaolong Su, Yujin Han, Feng Wang, Jameson Dong, Hins Hu, Difan Zou

CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models

Read the original on arXiv AI →

CoRe introduces a co‑evolving reward framework to mitigate latent reward hacking in video diffusion models. By continuously refitting the latent‑reward model on the generator’s current samples and anchoring it to real‑video preferences, CoRe prevents the generator from drifting outside the reward model’s training support. Experiments on Wan2.1‑T2V‑1.3B demonstrate that CoRe improves generation quality over pretrained models and prior alignment methods while avoiding quality collapse.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 25

Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation

arXiv:2608.21425v1 Announce Type: cross Abstract: Video generation is central to AI-powered content creation. Aligning generated videos with human preferences is a key criterion for evaluating genera...

By Nai-Xin Zhai, Weihua Cheng, Dexu Yu, Yikai Gu, Hanwen Du, Junchen Fu, Chenxi Huang, Yingwei Song, Liyuan Lillian Ma, Yang Ran, Youhua Li, Yongxin Ni
arXiv Machine Learning
Aug 5

Latent Reward Registers for Diffusion Preference Alignment

arXiv:2608. 03929v1 Announce Type: new Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process.

By Yuanshen Guan, Zipeng Feng, Zhiwei Xiong, Peiqin Sun
arXiv Computer Vision
Sep 7

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation

The paper introduces Objective-aware Trajectory Credit Assignment (OTCA), a framework that refines reinforcement learning for diffusion-based visual generation. OTCA decomposes credit across denoising steps and allocates multiple reward signals adaptively, addressing the coarse, uniform reward assignment of existing GRPO pipelines. Experiments demonstrate that OTCA consistently enhances image and video generation quality across various metrics.

By Rui Li, Ke Hao, Yuanzhi Liang, Haibin Huang, Chi Zhang, Yun Gu, Xuelong Li