arXiv AI

CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs

arXiv:2608. 14791v1 Announce Type: new Abstract: Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday sense-making and investigation to scientific discovery.

arXiv AI
Aug 18

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

arXiv:2604. 06628v2 Announce Type: replace Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes.

By Qihan Ren, Peng Wang, Ruikun Cai, Shuai Shao, Dadi Guo, Yuejin Xie, Yafu Li, Quanshi Zhang, Xia Hu, Jing Shao, Dongrui Liu
arXiv AI
Aug 12

UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention

arXiv:2607. 17188v2 Announce Type: replace Abstract: While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation, it can also exacerbate overthinking and underthinking, which we formulate as reasoning state--action mismatch.

By Cheng Yan, Zhijun Fan, Guangyang Ye, Fan Xu, Xiang Xia, Yawei Wang, Wuyang Zhang
arXiv AI
Sep 3

Thinking effort aligns between humans and reasoning models in abductive reasoning

The study examines how the effort expended by large reasoning models (LRMs) compares to that of humans during abductive reasoning tasks. By analyzing reaction times and reasoning traces, the authors find that LRMs and humans exhibit similar patterns of effort and error types. They also demonstrate that decoding strategies allowing models to explore multiple reasoning paths further align the models’ reasoning costs with human effort.

By Henry Arthur
arXiv Machine Learning
Sep 24

Giving Credit Where It's Due: Redundancy-Aware Learning for Efficient Reasoning

The paper introduces RECAP, a redundancy-aware credit assignment method that improves reasoning efficiency in large language models by assigning credit to each reasoning step based on its downstream role and contribution to the correct answer. RECAP uses a semantic dependency graph to measure structural responsibility and evaluates step efficacy via changes in gold-answer log-likelihood, enabling step-specific updates without requiring a separate reward model or concise trajectories. Experiments on two 7B models across four mathematical reasoning benchmarks show that RECAP enhances the accuracy-efficiency trade-off, boosting pass@1 by 2.0–3.7 percentage points while cutting reasoning tokens by 8–31% compared to GRPO.

By Yuqing Zhou, Hong Wang, Manqing Mao, Zhuoer Wang, Samson Koelle, Jie Yuan, Yanjun Lin, James Feng, Nikki Lijing Kuang, Ziwei Zhu, Wei Niu
arXiv AI
Jun 4

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots

arXiv:2606. 04503v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset.

By Guangcheng Zhu, Shenzhi Yang, Haobo Wang, Xing Zheng, Yingfan MA, Xuening Feng, Zhongqi Chen, Bowen Song, Weiqiang Wang, Gang Chen
arXiv AI
Jun 9

Correct Is Not Enough: Training Reasoning Planners with Executor-Grounded Rewards

arXiv:2605. 03862v4 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards has become a common way to improve explicit reasoning in large language models, but final-answer correctness alone does not reveal whether the reasoning trace is faithful, reliable, or useful to the model that consumes it.

By Tianyang Han, Hengyu Shi, Junjie Hu, Xu Yang, Zhiling Wang, Junhao Su
arXiv Machine Learning
Jun 5

SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions

arXiv:2604. 08477v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has substantially improved reasoning in formal domains such as mathematics and code, but extending these gains beyond STEM remains challenging.

By Ashima Suvarna, Kendrick Phan, Mehrab Beikzadeh, Hritik Bansal, Saadia Gabriel
arXiv AI
Sep 10

Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks

The paper investigates reinforcement learning with verifiable rewards (RLVR) by expanding the reasoning space beyond depth to include environment complexity and diverse reasoning forms. It introduces a synthetic knowledge‑graph environment that varies depth, complexity, and task family, revealing that joint depth‑complexity coverage outperforms single‑axis approaches, that different reasoning families behave non‑uniformly, and that uniform mixing beats staged curricula under a fixed budget. The study also shows that current off‑the‑shelf models share a deductive‑over‑abductive bias, indicating a broader gap in reasoning capabilities.

By Yihua Zhu, Qianying Liu, Fei Cheng, Jiaxin Wang, Akiko Aizawa, Sadao Kurohashi, Hidetoshi Shimodaira