arXiv AI By Yinghui Chi, Lucien Wang

Controllable and Verifiable Process Data Synthesis for Process Reward Models

Read the original on arXiv AI →

arXiv:2605. 02395v2 Announce Type: replace Abstract: Process reward models (PRMs) rely on high-quality process supervision data, yet existing construction methods often provide limited control over error location, error type, and trajectory consistency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

Learning Process Rewards via Reasoning State Propagation

The paper introduces Reasoning State Propagation (RSP), a method that models each reasoning prefix with a binary validity state and learns transitions between successive states. RSP predicts break and repair probabilities to connect intermediate reasoning states to the final outcome, enabling outcome supervision to guide learning of earlier steps. Experiments on reasoning search, response selection, and reinforcement learning show RSP consistently outperforms existing Process Reward Models, achieving notable gains over Qwen2.5-Math-PRM.

By Kai Gan, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei
arXiv Computation and Language
4d ago

RAWR: Reward Assignment Without Rollouts in Verifiable Domains

arXiv:2603.17815v2 Announce Type: replace Abstract: Understanding and evaluating multi-step reasoning in LLMs at the level of individual steps remains a key challenge. Process reward models (PRMs) pr...

By Corentin Royer (International Business Machines), Anna Hedstr\"om (ETH AI Center), Debarun Bhattacharjya (Lirio), Gaetano Rossiello (International Business Machines), Andrea Giovannini (International Business Machines), Mennatallah El-Assady (Department of Computer Science, ETH Zurich)
arXiv Computation and Language
Aug 28

Neuro-symbolic PRM: Enhancing Scientific Reasoning via Structured Traces and Symbolic Verification

The paper introduces a neuro‑symbolic framework for scientific reasoning that separates symbolic validity and semantic groundedness. A deterministic symbolic verifier acts as a hard filter to guarantee syntactic and arithmetic correctness, while a Process Reward Model (PRM) is trained on verifier‑accepted steps to assess contextual grounding. The authors propose Counterfactual Symbolic Perturbation (CSP) to generate hard negative examples that pass the verifier but are logically flawed, enabling efficient PRM training and a verifier‑first constrained search at inference.

By Yuxin Zi, Cong Xu, Suparna Bhattacharya, Martin Foltin, Amit Sheth