arXiv AI By Tianyang Han, Hengyu Shi, Junjie Hu, Xu Yang, Zhiling Wang, Junhao Su

Correct Is Not Enough: Training Reasoning Planners with Executor-Grounded Rewards

Read the original on arXiv AI →

arXiv:2605. 03862v4 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards has become a common way to improve explicit reasoning in large language models, but final-answer correctness alone does not reveal whether the reasoning trace is faithful, reliable, or useful to the model that consumes it.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 24

Giving Credit Where It's Due: Redundancy-Aware Learning for Efficient Reasoning

The paper introduces RECAP, a redundancy-aware credit assignment method that improves reasoning efficiency in large language models by assigning credit to each reasoning step based on its downstream role and contribution to the correct answer. RECAP uses a semantic dependency graph to measure structural responsibility and evaluates step efficacy via changes in gold-answer log-likelihood, enabling step-specific updates without requiring a separate reward model or concise trajectories. Experiments on two 7B models across four mathematical reasoning benchmarks show that RECAP enhances the accuracy-efficiency trade-off, boosting pass@1 by 2.0–3.7 percentage points while cutting reasoning tokens by 8–31% compared to GRPO.

By Yuqing Zhou, Hong Wang, Manqing Mao, Zhuoer Wang, Samson Koelle, Jie Yuan, Yanjun Lin, James Feng, Nikki Lijing Kuang, Ziwei Zhu, Wei Niu
arXiv Machine Learning
Aug 28

Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum

The paper investigates whether the high costs of training chain-of-thought reasoning models can be reduced through algorithmic design. It introduces an autocurriculum approach that lets the model select which problems to focus on during training, showing that this method provably improves both supervised fine‑tuning and reinforcement learning. For supervised fine‑tuning, autocurriculum requires exponentially fewer reasoning demonstrations by targeting prompts where the model struggles, while for reinforcement learning it decouples computational cost from the quality of the reference model, making the burn‑in cost nearly independent of target accuracy.

By Nived Rajaraman, Audrey Huang, Miro Dudik, Robert Schapire, Dylan J. Foster, Akshay Krishnamurthy