The paper introduces Reasoning State Propagation (RSP), a method that models each reasoning prefix with a binary validity state and learns transitions between successive states. RSP predicts break and repair probabilities to connect intermediate reasoning states to the final outcome, enabling outcome supervision to guide learning of earlier steps. Experiments on reasoning search, response selection, and reinforcement learning show RSP consistently outperforms existing Process Reward Models, achieving notable gains over Qwen2.5-Math-PRM.
By Kai Gan, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei
arXiv:2603.17815v2 Announce Type: replace
Abstract: Understanding and evaluating multi-step reasoning in LLMs at the level of individual steps remains a key challenge. Process reward models (PRMs) pr...
By Corentin Royer (International Business Machines), Anna Hedstr\"om (ETH AI Center), Debarun Bhattacharjya (Lirio), Gaetano Rossiello (International Business Machines), Andrea Giovannini (International Business Machines), Mennatallah El-Assady (Department of Computer Science, ETH Zurich)
arXiv:2512.03244v2 Announce Type: replace-cross
Abstract: Training process reward models (PRMs) requires step-level correctness labels, obtained either through expensive human annotation or by relyin...
By Salman Rahman, Sruthi Gorantla, Arpit Gupta, Swastik Roy, Nanyun Peng, Yang Liu
arXiv:2609.22746v1 Announce Type: new
Abstract: Large Language Models (LLMs) have recently been introduced into traffic signal control (TSC) as decision agents due to their strengths in human-readabl...
By Huaitao Zhao, Tianlong Zhou, Weijie Wang, Jiasheng Shi, Weixiong Rao
arXiv:2609.21492v1 Announce Type: new
Abstract: Chain-of-Thought (CoT) reasoning has been shown to improve the performance of large language models (LLMs), yet existing optimization methods largely r...
By Jingyu Hu, Shu Yang, Weiru Liu, Di Wang
The paper introduces a neuro‑symbolic framework for scientific reasoning that separates symbolic validity and semantic groundedness. A deterministic symbolic verifier acts as a hard filter to guarantee syntactic and arithmetic correctness, while a Process Reward Model (PRM) is trained on verifier‑accepted steps to assess contextual grounding. The authors propose Counterfactual Symbolic Perturbation (CSP) to generate hard negative examples that pass the verifier but are logically flawed, enabling efficient PRM training and a verifier‑first constrained search at inference.
By Yuxin Zi, Cong Xu, Suparna Bhattacharya, Martin Foltin, Amit Sheth