arXiv Machine Learning
Oct 2

Learning Infinite-Horizon Average-Reward CMDPs via State Augmentation

The paper introduces a computationally efficient algorithm for infinite-horizon average-reward constrained Markov decision processes (CMDPs) under weak communication. It achieves a high-probability regret and cumulative constraint violation of ×O(√T) in the tabular setting, matching optimal dependence up to logarithmic factors. The method augments the state with cumulative constraint violation, reshapes rewards using a Huber potential, and applies finite-horizon approximation with optimistic value iteration to maintain bounded per-step rewards.

By Kihyun Yu, Seoungbin Bae, Dabeen Lee
arXiv AI
Sep 15

ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement

ReCAST is a method for assigning credit to multiple rewards during diffusion model training by using a reward-by-timestep weight matrix that respects user-specified reward budgets while ensuring equal total weight per denoising step. It allocates weight based on each reward’s informativeness, measured by its Rényi discriminability gain at each step, allowing rewards to contribute more when they are most informative. Experiments on SD3.5‑Medium with two four‑reward settings show that ReCAST improves or matches training rewards, enhances held‑out judges, and is preferred by an independent LLM‑as‑a‑Judge, indicating generalizable benefits.

By Yihang Chen, Yuanhao Ban, Kuei-Chun Kao, Cho-Jui Hsieh