arXiv AI

ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement

ReCAST is a method for assigning credit to multiple rewards during diffusion model training by using a reward-by-timestep weight matrix that respects user-specified reward budgets while ensuring equal total weight per denoising step. It allocates weight based on each reward’s informativeness, measured by its Rényi discriminability gain at each step, allowing rewards to contribute more when they are most informative. Experiments on SD3.5‑Medium with two four‑reward settings show that ReCAST improves or matches training rewards, enhances held‑out judges, and is preferred by an independent LLM‑as‑a‑Judge, indicating generalizable benefits.

arXiv AI
4d ago

Diffusion Reward Models

arXiv:2609.33803v2 Announce Type: replace-cross Abstract: Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate...

By Xiangyang Wang, Bingxiang He, Zeyuan Liu, Jiaze Wang, Ziqing Qiao, Yuxin Zuo, Huan-ang Gao, Cheng Qian, Wenbin Zhang, Ran Li, Youbang Sun, Ning Ding, Yuanchun Shi, Zhiyuan Liu, Chaojun Xiao, Chun Yu
arXiv Machine Learning
Aug 5

Latent Reward Registers for Diffusion Preference Alignment

arXiv:2608. 03929v1 Announce Type: new Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process.

By Yuanshen Guan, Zipeng Feng, Zhiwei Xiong, Peiqin Sun
arXiv AI
Sep 24

Reinforcement Learning with Decomposed Subtasks

The paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS), a method that splits trajectory rewards into per‑subtask shares before policy updates, replacing the scalar advantage used in Group Relative Policy Optimization (GRPO). RLDS employs Subtask‑Decomposed Advantage Estimation (SDAE) to compute group‑relative advantages and distribute credit to tokens based on subtask importance, focusing on steps where a reflection marks a subtask as consequential. Experiments on four benchmarks—FrozenLake, HotpotQA, ScienceWorld, and DeepResearch—show that RLDS improves performance on high‑heterogeneity tasks (ScienceWorld and FrozenLake) and is more compute‑efficient than scalar GRPO for long rollouts.

By Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich
arXiv Machine Learning
Jun 18

The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL

arXiv:2606. 19162v1 Announce Type: new Abstract: Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with subjective preferences and, surprisingly, recovering properties such as visual realism and coherent object structure that matching-based training is intended to learn from the data itself.

By Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng, Reyhane Askari-Hemmat, Xiaochuang Han, Adriana Romero-Soriano, Michal Drozdzal
arXiv AI
Sep 24

When and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment

The paper introduces UECR-GRPO, a method that unifies on‑policy distillation and verifier‑based reinforcement learning for mathematical reasoning. It combines verifier rewards and teacher‑derived log‑ratios into a single KL‑regularized objective (Path‑Utility Unification) and then redistributes credit at the token level using entropy‑calibrated redistribution, preserving total task credit. Experiments on five benchmarks show that UECR‑GRPO improves average accuracy by up to 0.89 percentage points over the best baseline for both Qwen3‑1.7B and Qwen3‑4B students.

By Jie Zhang, Jingxiao Yang, Zhehao Huang, Yuhang Liu, Xiaolin Huang