arXiv:2606. 11521v1 Announce Type: new Abstract: LLMs and LLM agents should improve when given feedback, but identifying when they are able to do so is difficult: feedback is heterogeneous, domain-specific, and difficult to control.
By Hongyi Liu, Frederic Sala, Thomas Reps, Adithya Murali
arXiv:2608. 08786v1 Announce Type: new Abstract: Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when final answers are correct.
By Wenyao Cui, Huaping Zhang, Yongyi Huang, Qiuchi Li, Jian Xu, Cheng-Lin Liu, Chunxiao Gao, Juan Wang, Baohua Zhang
arXiv:2608. 19009v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly paired with verifiers (step checkers, self-consistency filters, tool-based fact checkers, formal proof assistants) that claim to detect the model's errors.
By Yajie Yin
The paper introduces a neuro‑symbolic framework for scientific reasoning that separates symbolic validity and semantic groundedness. A deterministic symbolic verifier acts as a hard filter to guarantee syntactic and arithmetic correctness, while a Process Reward Model (PRM) is trained on verifier‑accepted steps to assess contextual grounding. The authors propose Counterfactual Symbolic Perturbation (CSP) to generate hard negative examples that pass the verifier but are logically flawed, enabling efficient PRM training and a verifier‑first constrained search at inference.
By Yuxin Zi, Cong Xu, Suparna Bhattacharya, Martin Foltin, Amit Sheth
arXiv:2607. 23019v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting enables large language models (LLMs) to tackle multi-step reasoning tasks, yet the generated intermediate steps are not guaranteed to be logically sound.
By Zirong Chen, Meiyi Ma
arXiv:2607. 04572v1 Announce Type: new Abstract: Large language model (LLM) tutors often produce fluent step-by-step explanations, but a correct and pedagogically formatted response does not guarantee that the answer was derived from the student-facing problem.
By Bonan Shen, Dingyan Shang, Youting Wang, Tao Ning
MathAdv is a diagnostic benchmark for formal theorem proving that covers 13 undergraduate- and graduate-level mathematics domains. It includes Lean 4 proofs and up to three auxiliary tasks—multiple-choice questions, fill-in-the-blank problems, and expert-crafted transformations—to probe knowledge, informal reasoning, and robustness to problem presentation. Evaluation of current theorem provers shows formalization is a major bottleneck, performance varies by domain, natural-language guidance can help or hinder models, and equivalent reformulations reveal significant robustness gaps.
By Jiaxin Yuan, Connor Martinez Lockhart, Xiaoyu Liu, Jiaqi Wang, Chenghao Deng, Xiayimei Han, Vlasios Mastrantonis, Dmitrii Gudin, Shaopeng Zhu, Abdirisak Abdullahi Mohamed, Bilal Hamdi Aytekin, Jiewen Lang, Zezheng Song, Furong Huang
The paper introduces Evidence‑Diagnosed Intervention Training (EDIT), a two‑phase framework designed to improve rubric‑faithful grading by large language models. EDIT‑SFT first identifies problematic reasoning steps using internal model signals—posterior belief over the final mark and input‑grounding scores—and revises only those steps with rubric checklists. EDIT‑RL then calibrates the grader with belief‑guided reward shaping, penalising harmful belief drifts while encouraging useful exploration. Experiments on two real‑world, multi‑subject grading benchmarks show that EDIT consistently outperforms strong supervised fine‑tuning and reinforcement learning baselines, with ablation studies confirming the importance of internal‑state diagnostics.
By Zhihao Wu, Linhai Zhang, Taiyi Wang, Runcong Zhao, Peter Andrews, Cesare Aloisi, Yulan He
The paper introduces a verifier‑guided explainable reasoning framework for educational question answering that integrates gold‑anchored QLoRA, a task‑aware symbolic router, and group‑relative RLVR. It adapts Qwen2.5‑3B‑Instruct with field‑weighted QLoRA supervision, routes logic problems to a FOL/Z3 verifier and physics problems to a symbolic solver, and uses verifier feedback for candidate evaluation, self‑revision, and reward construction. Experiments on 438 held‑out examples show that RLVR boosts reasoning depth (P3) from 50.68 % to 72.20 %, while symbolic verification improves answer reliability at the system level.
By Thi Kim Trang Vo, Nam Tien Le, Thi Kim Nguyet Vo, Minh Khang Tran, Duy Phuong Tran
arXiv:2606. 15258v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly capable of mathematical problem solving and can even assist with research-level proofs, yet we still lack a scalable and reproducible way to measure step-level reasoning in long proofs across diverse sources.
By Jierui Zhang, Siyuan Tan, Xinhang Li, Longzhuangzhi Lin, Dailin Li, Chengfeng Gu, Xinping Li, Yaxian Hao, Shengjia Liang, Yuxiang Ren, Wenhao Liu
arXiv:2609.38409v1 Announce Type: new
Abstract: Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable r...
By \.Ibrahim Ethem Deveci, Funda Tan \c{C}al{\i}k, Bar{\i}\c{s} Deniz Sa\u{g}lam, Duygu Ataman
arXiv:2605. 23965v2 Announce Type: replace Abstract: Large Language Models (LLMs) achieve strong performance on logical reasoning benchmarks, yet their reliability remains uncertain.
By Zenghui Zhou, Man Li, Xiaoke Fang, Xinyi Zhou, Weibin Lin, Zheng Zheng