arXiv AI

Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs

arXiv:2510. 04140v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a widely adopted technique for enhancing the reasoning ability of Large Language Models (LLMs).

arXiv Computation and Language
Aug 28

Boosting LLM Exploration via Weak-Model Guidance in RLVR

The paper introduces a method to enhance large language model (LLM) exploration in Reinforcement Learning with Verifiable Rewards (RLVR) by guiding the target model with partial reasoning trajectories from smaller, weaker language models. This weak-model guidance disrupts over‑confidence, preserves generative diversity, and mitigates entropy collapse without extra fine‑tuning or complex reward designs. Experiments on mathematical benchmarks show consistent improvements over vanilla RLVR, especially as the number of allowed attempts ($k$) increases, indicating broader reasoning coverage.

By Xingyu Shen, Huishuai Zhang, Peng Li, Yinchun Wang, Dongyan Zhao
arXiv AI
Sep 10

Boosting LLM Reasoning via Human-Inspired Reward Shaping

The paper introduces T2T (Thickening-to-Thinning), a dynamic reward framework for large language models that mimics human learning by separating exploration and consolidation phases. During incorrect attempts, T2T encourages exploration to broaden the search space, while after correct solutions it applies length penalties to promote concise reasoning. Experiments on mathematical benchmarks across five mainstream LLMs show that T2T outperforms standard GRPO and recent baselines, improving overall reasoning performance.

By Wenze Lin, Zhen Yang, Xitai Jiang, Xiaoteng Ma, Gao Huang
arXiv Machine Learning
Jun 9

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning

arXiv:2606. 08088v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has recently become a key paradigm for improving the reasoning abilities of Large Language Models (LLMs), yet it remains limited by sparse binary rewards and its ignorance of model-internal uncertainty.

By Qing Miao, Yiming Zhao, Jing Yang, Chenxi Liu, Yuehai Chen, Yuewen Liu, Shaoyi Du, Badong Chen
Hugging Face Trending Papers
Jun 24

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources

Achieving strong optimization generalization across diverse optimization problems while requiring limited training resources remains a challenging problem for optimization-oriented large language models (LLMs). Existing approaches typically rely on large-scale supervised datasets, costly reasoning annotations, and expensive intermediate step verification, resulting in substantial training overhead.

arXiv AI
Jun 24

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

arXiv:2606. 24064v1 Announce Type: new Abstract: Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather than how to reason.

By Tianyuan Shi, Canbin Huang, Bei Li, Xin Chen, Xiaojun Quan, Jingang Wang, Qifan Wang
arXiv AI
2d ago

Improving Math Reasoning through Value-guided Informative Search

The paper introduces APIVIS, a training-time framework that integrates finite-budget Gumbel search into reinforcement learning with verifiable rewards (RLVR) for mathematical reasoning. APIVIS combines direct and searched responses within each rollout group, uses selective supervision on search-improved tokens, and applies value-guided selection to improve verifier rewards at each searched state. Experiments on standard mathematical reasoning benchmarks and various model scales show significant performance gains over existing search-based methods.

By Shaohuai Liu, Yuning Wu, Haoran Liu, Enzo Jia, Devin Chen, Kai Wei
arXiv AI
Sep 10

Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR

The paper introduces DATPO, a Difficulty‑Adaptive Sentence‑entropy‑guided Tree‑structured Policy Optimization method designed to improve reasoning coverage in Reinforcement Learning with Verifiable Rewards (RLVR). It builds on three design principles: adaptive difficulty rollouts, tree‑based rollouts, and sentence‑entropy‑guided forking to enhance semantic diversity. Experiments on mathematical reasoning benchmarks show that DATPO outperforms existing baselines, particularly in pass@k, leading to better test‑time scaling performance.

By Youngjun Yu, Sanghwan Jang, Hwanjo Yu