arXiv AI

MCTS-KBQA: Monte Carlo Tree Search with Information Gain Rewards for Knowledge Base Question Answering

The paper introduces Fast MCTS, a Monte Carlo Tree Search approach for knowledge base question answering that replaces costly terminal rollouts with an information gain reward for intermediate states. This reward is computed using a question‑conditioned PPL‑ratio proxy over sanitized interaction histories, leveraging an open‑source instruction LLM without extra training. Experiments on four KBQA benchmarks demonstrate that Fast MCTS consistently outperforms linear baselines and improves the accuracy‑cost trade‑off compared to classic rollout‑based MCTS.

arXiv AI
3d ago

Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning

Goldilocks RL proposes an adaptive data‑selection strategy for reinforcement learning in language models, using a Selector network to predict reward variability and prioritize questions that are neither too easy nor too hard. By continuously adapting to the model’s evolving abilities, Goldilocks improves training efficiency on large‑scale reasoning datasets, achieving up to 78% fewer optimization steps compared to standard GRPO. The approach addresses the sample‑inefficiency problem of sparse rewards in reasoning tasks.

By Ilia Mahrooghi, Aryo Lotfi, Emmanuel Abbe
arXiv Machine Learning
Sep 14

Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models

The paper introduces Decision-Flow Sampling (DF‑Sample), a training‑free, data‑free inference framework that builds a hierarchical reasoning tree, evaluates entire trajectories, and back‑propagates utilities to guide branching decisions. Unlike local step‑wise sampling, DF‑Sample explicitly assesses global paths, enabling it to recover high‑quality, low‑probability reasoning chains that standard decoding misses. On the GPQA benchmark, DF‑Sample attains 45.6% accuracy, outperforming power sampling (38.9%) and GRPO (39.9%) and consistently surpassing baselines across multiple models and benchmarks, demonstrating significant latent reasoning potential in pretrained LLMs.

By Zhendong Mi, Shaoyi Huang
arXiv Machine Learning
Aug 24

Reinforcing Multi-Turn Reasoning in LLM Agents via Fine-Grained Reward Structure and Credit Assignment

The paper explores how dense, turn-level reward structures can improve reinforcement learning for large language model agents in multi-turn tasks. It introduces three reward granularity types—terminal, delayed, and per-turn—and adapts Group Relative Policy Optimization and Proximal Policy Optimization to each. Experiments on search and game agents show that per-turn rewards consistently yield better training dynamics, faster convergence, and higher answer correctness compared to sparse terminal or delayed rewards.

By Quan Wei, Siliang Zeng, Chenliang Li, Zhongruo Wang, William Brown, Oana Frunza, Wei Deng, Anderson Schneider, Yuriy Nevmyvaka, Yang Katie Zhao, Alfredo Garcia, Mingyi Hong
arXiv AI
Oct 1

StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning

StateTree is a reinforcement learning approach that improves long‑term dialogue reasoning by building a tree‑structured auxiliary task from limited dialogue data. The method embeds key‑value records across multiple sessions into a binary tree, requiring the model to traverse from root to leaf, retrieve records, compare timestamps, and identify a target question among distractors. Curriculum RL training increases tree depth, and a compositional variant trains the model to combine partial reasoning fragments, enabling cross‑session retrieval, temporal reasoning, knowledge updates, and multi‑hop reasoning while generalizing from 10K‑token to 128K‑token contexts.

By Naen Xu, Wanqing Cui, Yibo Hu, Shixin Hong, Hengyu An, Meiguang Jin, Junfeng Ma, Tianyu Du