arXiv AI

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion

arXiv:2607. 20543v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling.

arXiv AI
Sep 1

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

The paper investigates why reinforcement learning with verifiable rewards (RLVR) reduces the diversity of solutions in reasoning tasks. By analyzing the Countdown task, the authors show that RLVR contracts the solution space mainly at the entrance—before the first arithmetic operation—causing a 67% drop in solution coverage. They demonstrate that providing an unselected entrance prefix or applying entrance‑targeted interventions can restore or even improve coverage without harming accuracy.

By Qiancheng Zhou, Ruizhe Li
arXiv AI
Jun 18

Sparsity Curse: Understanding RLVR Model Parameter Space from Model Merging

arXiv:2606. 18521v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Reward (RLVR) has emerged as a powerful post-training paradigm that surpasses Supervised Fine-Tuning (SFT) in eliciting reasoning intelligence and resisting catastrophic forgetting.

By Chenrui Wu, Zexi Li, Jiajun Bu, Jiangchuan Liu, Haishuai Wang
arXiv Machine Learning
Sep 22

RLVR is a Kernel, Not a Function: Statistical Inference for pass@$k$ Crossovers

The paper argues that the observed crossover in pass@$k$ performance between reinforcement learning with verifiable rewards (RLVR) and its base model is not always statistically confirmed. By constructing confidence bands across sampling budgets and performing power analyses, the authors show that many reported crossovers lack statistical support and that more prompts can improve detection more than more answers per prompt. They further demonstrate that RLVR’s effect on a prompt is better described as a conditional distribution (a Markov kernel) rather than a single curve, allowing predictions of crossovers in new data and clarifying how losses on difficult prompts can overturn early gains.

By Chen Yang, Xianyang Zhang, Jun Chen