arXiv Computation and Language By Yongqi Tong, Zhenyu Zhang, Zimi Liu, Kewei Fu, Mingli Song, Haofei Zhang, Junshao Zhang, Hong Zhu, Jiang-Ming Yang, Xin Zhang, Jianshe Li

Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

Read the original on arXiv Computation and Language →

The paper introduces Ask-Condition-Abstain Reinforcement Learning (ACA‑RL), a framework that trains reasoning models to handle queries missing a premise by either asking for it, conditioning on the unknown, or abstaining. ACA‑RL uses a reasoning‑graph‑guided pipeline to generate training instances with localized gap annotations and a structured reward over five observable response behaviors. The authors also present the Missing‑Premise Benchmark (MPB), a 274‑instance, human‑verified dataset covering mathematical, logical, and real‑world word problems, and show that ACA‑RL improves performance on MPB while maintaining competitive results on well‑posed tasks for Qwen3 and Llama models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
4d ago

Learning to Prove, Not Just to Answer: Reinforcement Learning from Formal Verification for Natural-Language Logical Reasoning

The paper introduces Proof‑R1, a reinforcement‑learning framework that trains large language models to generate verifiable proofs for natural‑language logical reasoning tasks. Proof‑R1 only accepts a generated conclusion into the proof state when it satisfies formal verification constraints, ensuring each reasoning step is machine‑checkable. The method also reconstructs the dependency closure that supports the final answer, aligning credit with valid proof steps, and shows improved answer accuracy and verifiability across multiple benchmarks and models.

By Qili Zhang, Qianren Mao, Hanze Cai, Kaiming Zhao, Yuening He, Xihan Lei, Yashuo Luo, Hanwen Hao, Yutong Gu, Likang Xiao, Zhijun Chen, Weifeng Jiang, Haoyi Zhou, Jianxin Li
Hugging Face Trending Papers
Jun 1

Learning When to Translate for Multilingual Reasoning

Reasoning language models (RLMs) achieve strong performance on complex reasoning tasks, but still exhibit substantial multilingual reasoning gaps, largely due to language-understanding failures in non-English inputs. English translation can mitigate these failures by expressing non-English inputs in a form that RLMs can more reliably interpret, yet translating every input is unnecessary when the model can reason reliably from the original query.