arXiv AI

SymStep: Symbolic Step Verification for Logical Reasoning

arXiv:2607. 23055v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently across steps.

arXiv Computation and Language
Aug 28

Neuro-symbolic PRM: Enhancing Scientific Reasoning via Structured Traces and Symbolic Verification

The paper introduces a neuro‑symbolic framework for scientific reasoning that separates symbolic validity and semantic groundedness. A deterministic symbolic verifier acts as a hard filter to guarantee syntactic and arithmetic correctness, while a Process Reward Model (PRM) is trained on verifier‑accepted steps to assess contextual grounding. The authors propose Counterfactual Symbolic Perturbation (CSP) to generate hard negative examples that pass the verifier but are logically flawed, enabling efficient PRM training and a verifier‑first constrained search at inference.

By Yuxin Zi, Cong Xu, Suparna Bhattacharya, Martin Foltin, Amit Sheth
arXiv AI
2d ago

The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning

The paper introduces a new approach to distill reasoning abilities from large language models (LLMs) into smaller student models by framing the task as a constrained reinforcement learning problem. It enforces a worst‑case constraint on the teacher’s log‑likelihood for every prefix of the reasoning chain, avoiding reward hacking and excessive teacher regularization. Experiments on mathematical reasoning and code generation show that this method improves the balance between accuracy and fidelity, achieving the highest rigorous reasoning success rate among evaluated settings.

By Matthieu Zimmer, Xiaotong Ji, Tu Nguyen, Haitham Bou-Ammar