arXiv AI
3d ago

Learning Process Rewards via Reasoning State Propagation

The paper introduces Reasoning State Propagation (RSP), a method that models each reasoning prefix with a binary validity state and learns transitions between successive states. RSP predicts break and repair probabilities to connect intermediate reasoning states to the final outcome, enabling outcome supervision to guide learning of earlier steps. Experiments on reasoning search, response selection, and reinforcement learning show RSP consistently outperforms existing Process Reward Models, achieving notable gains over Qwen2.5-Math-PRM.

By Kai Gan, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei
arXiv Machine Learning
Sep 24

Giving Credit Where It's Due: Redundancy-Aware Learning for Efficient Reasoning

The paper introduces RECAP, a redundancy-aware credit assignment method that improves reasoning efficiency in large language models by assigning credit to each reasoning step based on its downstream role and contribution to the correct answer. RECAP uses a semantic dependency graph to measure structural responsibility and evaluates step efficacy via changes in gold-answer log-likelihood, enabling step-specific updates without requiring a separate reward model or concise trajectories. Experiments on two 7B models across four mathematical reasoning benchmarks show that RECAP enhances the accuracy-efficiency trade-off, boosting pass@1 by 2.0–3.7 percentage points while cutting reasoning tokens by 8–31% compared to GRPO.

By Yuqing Zhou, Hong Wang, Manqing Mao, Zhuoer Wang, Samson Koelle, Jie Yuan, Yanjun Lin, James Feng, Nikki Lijing Kuang, Ziwei Zhu, Wei Niu