arXiv AI By Kai Gan, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei

Learning Process Rewards via Reasoning State Propagation

Read the original on arXiv AI →

The paper introduces Reasoning State Propagation (RSP), a method that models each reasoning prefix with a binary validity state and learns transitions between successive states. RSP predicts break and repair probabilities to connect intermediate reasoning states to the final outcome, enabling outcome supervision to guide learning of earlier steps. Experiments on reasoning search, response selection, and reinforcement learning show RSP consistently outperforms existing Process Reward Models, achieving notable gains over Qwen2.5-Math-PRM.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 25

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization

The paper introduces the Implicit Prefix-Value Reward Model (IPVRM), which learns the probability of eventual correctness for each prefix directly from outcome labels, thereby aligning training targets with inference-time step signals via temporal-difference differences. IPVRM improves step-verification F1 on ProcessBench. Additionally, the authors propose Distribution-Level RL (DistRL), a policy optimization method that applies TD advantages to both sampled and high-probability tokens, offering dense counterfactual updates without extra rollouts, and show that DistRL consistently enhances downstream reasoning when combined with IPVRM.

By Shiping Gao, Hongzhan Chen, Xiaojun Quan, Qifan Wang, Lifu Huang