The paper introduces Reasoning State Propagation (RSP), a method that models each reasoning prefix with a binary validity state and learns transitions between successive states. RSP predicts break and repair probabilities to connect intermediate reasoning states to the final outcome, enabling outcome supervision to guide learning of earlier steps. Experiments on reasoning search, response selection, and reinforcement learning show RSP consistently outperforms existing Process Reward Models, achieving notable gains over Qwen2.5-Math-PRM.
By Kai Gan, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei
arXiv:2603.17815v2 Announce Type: replace
Abstract: Understanding and evaluating multi-step reasoning in LLMs at the level of individual steps remains a key challenge. Process reward models (PRMs) pr...
By Corentin Royer (International Business Machines), Anna Hedstr\"om (ETH AI Center), Debarun Bhattacharjya (Lirio), Gaetano Rossiello (International Business Machines), Andrea Giovannini (International Business Machines), Mennatallah El-Assady (Department of Computer Science, ETH Zurich)
arXiv:2512.03244v2 Announce Type: replace-cross
Abstract: Training process reward models (PRMs) requires step-level correctness labels, obtained either through expensive human annotation or by relyin...
By Salman Rahman, Sruthi Gorantla, Arpit Gupta, Swastik Roy, Nanyun Peng, Yang Liu
arXiv:2609.22746v1 Announce Type: new
Abstract: Large Language Models (LLMs) have recently been introduced into traffic signal control (TSC) as decision agents due to their strengths in human-readabl...
By Huaitao Zhao, Tianlong Zhou, Weijie Wang, Jiasheng Shi, Weixiong Rao
arXiv:2609.21492v1 Announce Type: new
Abstract: Chain-of-Thought (CoT) reasoning has been shown to improve the performance of large language models (LLMs), yet existing optimization methods largely r...
By Jingyu Hu, Shu Yang, Weiru Liu, Di Wang
The paper introduces a neuro‑symbolic framework for scientific reasoning that separates symbolic validity and semantic groundedness. A deterministic symbolic verifier acts as a hard filter to guarantee syntactic and arithmetic correctness, while a Process Reward Model (PRM) is trained on verifier‑accepted steps to assess contextual grounding. The authors propose Counterfactual Symbolic Perturbation (CSP) to generate hard negative examples that pass the verifier but are logically flawed, enabling efficient PRM training and a verifier‑first constrained search at inference.
By Yuxin Zi, Cong Xu, Suparna Bhattacharya, Martin Foltin, Amit Sheth
arXiv:2601. 22642v2 Announce Type: replace Abstract: Large Language Models (LLMs) show remarkable capabilities, yet their stochastic next-token prediction creates logical inconsistencies and reward hacking that formal symbolic systems avoid.
By Chuxue Cao, Jinluan Yang, Haoran Li, Kunhao Pan, Zijian Zhao, Zhengyu Chen, Yuchen Tian, Lijun Wu, Conghui He, Sirui Han, Yike Guo
arXiv:2605. 12519v2 Announce Type: replace-cross Abstract: Training language models to produce both correct answers and sound reasoning remains an open challenge.
By Kyuyoung Kim, Kevin Wang, Yunfei Xie, Peiyang Xu, Peiyao Sheng, Chen Wei, Zhangyang Wang, Jinwoo Shin, Pramod Viswanath, Sewoong Oh
arXiv:2505. 04671v3 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) trained with reinforcement learning (RL) have improved Text-to-SQL performance.
By Yuxin Zhang, Meihao Fan, Ju Fan, Mingyang Yi, Yuyu Luo, Guoliang Li, Bin Wu, Wenchao Zhou
The paper introduces MIMIC, a framework that uses executable code to generate rigorous reasoning data for large language models (LLMs). By converting algorithms into verifiable reasoning trajectories through narrative fusion, code-guided test synthesis, and dynamic code instrumentation, MIMIC creates a Code-Instrumented Reward (CIR) that supplies dense, high‑fidelity supervision for reinforcement learning. Models trained with MIMIC’s synthetic dataset show significant, consistent improvements in general reasoning, complex mathematics, and fine‑grained deterministic tasks.
By Jinyang Zhang, Weibin Liao, Keqin Bao, Sihang Li, Shaobo Wang, Muyang Ye, Hongxin Ding, Yue Fang, Tianyi Tang, Fei Huang, Kexin Yang, Xingzhang Ren, Dayiheng Liu
The paper introduces LSR‑Ben, a benchmark designed to evaluate process reward models (PRMs) on scientific and logical reasoning tasks, addressing a gap left by existing math‑focused benchmarks. Experiments on 22 models reveal that PRMs and LLMs perform poorly in non‑mathematical domains, with LLMs tending to over‑identify errors while PRMs tend to overlook them. LSR‑Ben aims to spur research that broadens PRM applicability and improves LLM reasoning.
By Zhouhao Sun, Xuan Zhang, Xiao Ding, Bibo Cai, Li Du, Kai Xiong, Xinran Dai, Fei Zhang, weidi tang, Zhiyuan Kan, Yang Zhao, Bing Qin, Ting Liu
arXiv:2604. 09482v2 Announce Type: replace Abstract: Reasoning in knowledge-intensive domains remains challenging as intermediate steps are often not locally verifiable: unlike math or code, evaluating step correctness may require synthesizing clues across large external knowledge sources.
By Jiwoong Sohn, Tomasz Sternal, Kenneth Styppa, Torsten Hoefler, Michael Moor