Scaling laws for reward model overoptimization
Read the original on OpenAI Blog →The Flow has not summarised this story yet — read it at OpenAI Blog.
The Flow has not summarised this story yet — read it at OpenAI Blog.
arXiv:2602. 02572v2 Announce Type: replace-cross Abstract: Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy.
arXiv:2607. 18300v1 Announce Type: cross Abstract: We extend Incentive Compatible Exploration beyond the Bayesian full-information setting of Kremer et al.
We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require c...
The paper introduces a computationally efficient algorithm for infinite-horizon average-reward constrained Markov decision processes (CMDPs) under weak communication. It achieves a high-probability regret and cumulative constraint violation of ×O(√T) in the tabular setting, matching optimal dependence up to logarithmic factors. The method augments the state with cumulative constraint violation, reshapes rewards using a Huber potential, and applies finite-horizon approximation with optimistic value iteration to maintain bounded per-step rewards.
arXiv:2601. 21523v2 Announce Type: replace Abstract: To promote cooperation in Multi-Agent Reinforcement Learning, the reward signals of all agents can be aggregated together, forming global rewards that are commonly known as the fully cooperative setting.
ReCAST is a method for assigning credit to multiple rewards during diffusion model training by using a reward-by-timestep weight matrix that respects user-specified reward budgets while ensuring equal total weight per denoising step. It allocates weight based on each reward’s informativeness, measured by its Rényi discriminability gain at each step, allowing rewards to contribute more when they are most informative. Experiments on SD3.5‑Medium with two four‑reward settings show that ReCAST improves or matches training rewards, enhances held‑out judges, and is preferred by an independent LLM‑as‑a‑Judge, indicating generalizable benefits.