arXiv AI By Bin Lei, Yu Li, Prafulla Kumar Choubey, Jiaxin Zhang, Becky Xiangyu Peng, Qinyuan Ye, Kartik Narayan, Caiwen Ding, Silvio Savarese, Chien-Sheng Wu

Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning

Read the original on arXiv AI →

The paper introduces belief‑shift branching, a method for placing forks in tree‑structured reinforcement learning rollouts by detecting where a model’s answer belief changes most sharply. Unlike traditional fixed‑length or entropy‑based forking, this approach uses a lightweight probe or learned activation direction to identify pivots in the value curve, reducing unnecessary sampling. Experiments show that belief‑shift forking consistently outperforms baseline methods across multiple models and benchmarks, yielding significant gains in mathematics and code tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 10

Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning

The paper introduces belief‑shift branching, a method for placing forks in tree‑structured reinforcement learning rollouts by identifying points where a model’s answer belief changes most. Unlike traditional structural or entropy‑based approaches, belief‑shift uses a probe, logit‑lens depth profile, or learned activation direction to locate pivots in the value curve, incurring minimal computational overhead. Experiments across multiple models and benchmarks show that belief‑shift forking consistently outperforms baseline methods, yielding significant gains in mathematics and code tasks.

arXiv AI
Sep 18

EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning

EPIG-Tree proposes a compute‑optimal branching strategy for gradient‑efficient reinforcement learning, arguing that branches should be placed where they most reduce policy‑gradient uncertainty per unit of compute. By deriving allocation laws from a law‑of‑total‑variance decomposition, the method introduces an EPIG‑Tree score that guides branch placement using already computed rollouts, estimating occupancy‑ and score‑weighted value uncertainty. Empirical results show EPIG‑Tree reduces gradient MSE in cloned‑state control, improves frozen‑LLM gradient calibration, and outperforms flat GRPO and entropy branching in both single‑turn math and multi‑turn Wordle tasks.

By Nikita Khomich, Leopold Hermansson, Ido Hakimi
arXiv AI
Sep 1

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

The paper investigates why reinforcement learning with verifiable rewards (RLVR) reduces the diversity of solutions in reasoning tasks. By analyzing the Countdown task, the authors show that RLVR contracts the solution space mainly at the entrance—before the first arithmetic operation—causing a 67% drop in solution coverage. They demonstrate that providing an unselected entrance prefix or applying entrance‑targeted interventions can restore or even improve coverage without harming accuracy.

By Qiancheng Zhou, Ruizhe Li