arXiv Machine Learning By Zihan Liu, Xurong Xie

Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning

Read the original on arXiv Machine Learning →

The paper introduces Function-Structured Graph Reinforcement Learning (FSG‑RL), a framework that links subproblem graphs to Python code and uses multiple verifiers for feedback. It first trains a policy via supervised fine‑tuning to generate code from function graphs, then refines it with Group Relative Policy Optimization (GRPO) that employs answer‑gated rewards and span‑level credit assignment. On a benchmark combining GSM8K, MathQA, MATH, and Omni‑MATH, GRPO raises final‑answer accuracy from 43.25 % to 67.50 % and full‑solution success from 32.25 % to 52.25 %, with further improvements when teacher supervision is added.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
4d ago

Learning to Prove, Not Just to Answer: Reinforcement Learning from Formal Verification for Natural-Language Logical Reasoning

The paper introduces Proof‑R1, a reinforcement‑learning framework that trains large language models to generate verifiable proofs for natural‑language logical reasoning tasks. Proof‑R1 only accepts a generated conclusion into the proof state when it satisfies formal verification constraints, ensuring each reasoning step is machine‑checkable. The method also reconstructs the dependency closure that supports the final answer, aligning credit with valid proof steps, and shows improved answer accuracy and verifiability across multiple benchmarks and models.

By Qili Zhang, Qianren Mao, Hanze Cai, Kaiming Zhao, Yuening He, Xihan Lei, Yashuo Luo, Hanwen Hao, Yutong Gu, Likang Xiao, Zhijun Chen, Weifeng Jiang, Haoyi Zhou, Jianxin Li
arXiv AI
Aug 28

From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning

The paper investigates the reliability of rule- and model-based verifiers used in reinforcement learning with verifiable reward (RLVR) for mathematical reasoning. It finds that rule-based verifiers often miss equivalent answers in different formats, causing false negatives that degrade RL performance as models improve. Model-based verifiers achieve higher static accuracy but become vulnerable to reward hacking during RL, misclassifying certain response patterns as correct after fine-tuning.

By Yuzhen Huang, Weihao Zeng, Xingshan Zeng, Qi Zhu, Junxian He
arXiv AI
2d ago

Improving Math Reasoning through Value-guided Informative Search

The paper introduces APIVIS, a training-time framework that integrates finite-budget Gumbel search into reinforcement learning with verifiable rewards (RLVR) for mathematical reasoning. APIVIS combines direct and searched responses within each rollout group, uses selective supervision on search-improved tokens, and applies value-guided selection to improve verifier rewards at each searched state. Experiments on standard mathematical reasoning benchmarks and various model scales show significant performance gains over existing search-based methods.

By Shaohuai Liu, Yuning Wu, Haoran Liu, Enzo Jia, Devin Chen, Kai Wei