arXiv Machine Learning

RLVR is a Kernel, Not a Function: Statistical Inference for pass@$k$ Crossovers

The paper argues that the observed crossover in pass@$k$ performance between reinforcement learning with verifiable rewards (RLVR) and its base model is not always statistically confirmed. By constructing confidence bands across sampling budgets and performing power analyses, the authors show that many reported crossovers lack statistical support and that more prompts can improve detection more than more answers per prompt. They further demonstrate that RLVR’s effect on a prompt is better described as a conditional distribution (a Markov kernel) rather than a single curve, allowing predictions of crossovers in new data and clarifying how losses on difficult prompts can overturn early gains.

Hugging Face Trending Papers
Sep 10

Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning

The paper introduces belief‑shift branching, a method for placing forks in tree‑structured reinforcement learning rollouts by identifying points where a model’s answer belief changes most. Unlike traditional structural or entropy‑based approaches, belief‑shift uses a probe, logit‑lens depth profile, or learned activation direction to locate pivots in the value curve, incurring minimal computational overhead. Experiments across multiple models and benchmarks show that belief‑shift forking consistently outperforms baseline methods, yielding significant gains in mathematics and code tasks.

arXiv AI
Sep 12

Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning

The paper introduces belief‑shift branching, a method for placing forks in tree‑structured reinforcement learning rollouts by detecting where a model’s answer belief changes most sharply. Unlike traditional fixed‑length or entropy‑based forking, this approach uses a lightweight probe or learned activation direction to identify pivots in the value curve, reducing unnecessary sampling. Experiments show that belief‑shift forking consistently outperforms baseline methods across multiple models and benchmarks, yielding significant gains in mathematics and code tasks.

By Bin Lei, Yu Li, Prafulla Kumar Choubey, Jiaxin Zhang, Becky Xiangyu Peng, Qinyuan Ye, Kartik Narayan, Caiwen Ding, Silvio Savarese, Chien-Sheng Wu
arXiv Machine Learning
Aug 27

Demystifying Reinforcement Learning Post-Training of Language Models

The paper "Demystifying Reinforcement Learning Post-Training of Language Models" investigates how reinforcement learning (RL) post‑training enhances large language models (LLMs) for tasks such as reasoning, math, and coding. By isolating RL components in a controlled setting, the authors analyze how the base model’s prior distribution, reward granularity, prompt diversity, and model scale influence outcomes, using policy entropy to compare pre‑training, supervised fine‑tuning (SFT), and RL stages. The study clarifies the role of spurious rewards, the importance of the base model’s probability mass on desired behaviors, and how these factors interact to determine post‑training success, offering a practical primer for NLP researchers. "whyItMatters":"The work provides a clearer understanding of RL post‑training mechanics, helping researchers and practitioners effectively apply RL to improve LLM capabilities."

By Donovan Clay, Saket Gollapudi, Sankar Harilal, Min Jang, Jacob Morrison, Sewoong Oh, Natasha Jaques