arXiv AI By Guanming Xiong, Haochen Li, Zonghong Dai, Liqiang Wen, Wen Zhao

MCTS-KBQA: Monte Carlo Tree Search with Information Gain Rewards for Knowledge Base Question Answering

Read the original on arXiv AI →

The paper introduces Fast MCTS, a Monte Carlo Tree Search approach for knowledge base question answering that replaces costly terminal rollouts with an information gain reward for intermediate states. This reward is computed using a question‑conditioned PPL‑ratio proxy over sanitized interaction histories, leveraging an open‑source instruction LLM without extra training. Experiments on four KBQA benchmarks demonstrate that Fast MCTS consistently outperforms linear baselines and improves the accuracy‑cost trade‑off compared to classic rollout‑based MCTS.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning

Goldilocks RL proposes an adaptive data‑selection strategy for reinforcement learning in language models, using a Selector network to predict reward variability and prioritize questions that are neither too easy nor too hard. By continuously adapting to the model’s evolving abilities, Goldilocks improves training efficiency on large‑scale reasoning datasets, achieving up to 78% fewer optimization steps compared to standard GRPO. The approach addresses the sample‑inefficiency problem of sparse rewards in reasoning tasks.

By Ilia Mahrooghi, Aryo Lotfi, Emmanuel Abbe
arXiv Machine Learning
Sep 14

Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models

The paper introduces Decision-Flow Sampling (DF‑Sample), a training‑free, data‑free inference framework that builds a hierarchical reasoning tree, evaluates entire trajectories, and back‑propagates utilities to guide branching decisions. Unlike local step‑wise sampling, DF‑Sample explicitly assesses global paths, enabling it to recover high‑quality, low‑probability reasoning chains that standard decoding misses. On the GPQA benchmark, DF‑Sample attains 45.6% accuracy, outperforming power sampling (38.9%) and GRPO (39.9%) and consistently surpassing baselines across multiple models and benchmarks, demonstrating significant latent reasoning potential in pretrained LLMs.

By Zhendong Mi, Shaoyi Huang
arXiv Machine Learning
Aug 24

Reinforcing Multi-Turn Reasoning in LLM Agents via Fine-Grained Reward Structure and Credit Assignment

The paper explores how dense, turn-level reward structures can improve reinforcement learning for large language model agents in multi-turn tasks. It introduces three reward granularity types—terminal, delayed, and per-turn—and adapts Group Relative Policy Optimization and Proximal Policy Optimization to each. Experiments on search and game agents show that per-turn rewards consistently yield better training dynamics, faster convergence, and higher answer correctness compared to sparse terminal or delayed rewards.

By Quan Wei, Siliang Zeng, Chenliang Li, Zhongruo Wang, William Brown, Oana Frunza, Wei Deng, Anderson Schneider, Yuriy Nevmyvaka, Yang Katie Zhao, Alfredo Garcia, Mingyi Hong