The paper investigates why reinforcement learning with verifiable rewards (RLVR) reduces the diversity of solutions in reasoning tasks. By analyzing the Countdown task, the authors show that RLVR contracts the solution space mainly at the entrance—before the first arithmetic operation—causing a 67% drop in solution coverage. They demonstrate that providing an unselected entrance prefix or applying entrance‑targeted interventions can restore or even improve coverage without harming accuracy.
By Qiancheng Zhou, Ruizhe Li
The study investigates whether post‑training methods—GRPO, SFT, and DPO—improve language models’ ability to follow prompt evidence that conflicts with memorized knowledge. By comparing nine training variants across different scales and families, the authors find that grounding gains are modest for GRPO, moderate for Conflict‑SFT, and near‑ceiling for DPO, but all largely rely on the same causal attention‑head set present in the starting checkpoint. Removing the starting‑model grounding direction suppresses these gains, while adding it back recovers a significant portion of DPO’s improvement, indicating that existing model machinery drives most of the observed gains.
By Prakhar Gupta, Vaibhav Gupta
arXiv:2607. 10203v1 Announce Type: cross Abstract: Adaptive-compute world models -- early-exit or mixture-of-depths predictors that spend variable depth per step -- assume depth buys better predictions and can be routed adaptively.
By Achyuthan Sivasankar
arXiv:2607. 21372v1 Announce Type: cross Abstract: Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios.
By Jingyuan Li, Xiaoyi Jiang, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
Reinforcement learning from verifiable rewards (RLVR) for mathematical reasoning suffers from a structural blind spot: on "cliff" prompts-those on which every sampled rollout in a group fails-the group-normalized advantage is identically zero, so GRPO produces no gradient on precisely the prompts at the frontier of the model's capability. We introduce LoRA Scaffolded Policy Optimization (LSPO), a sampling-time mechanism that recovers this lost gradient.
arXiv:2609.09075v1 Announce Type: cross
Abstract: In reinforcement learning with verifiable rewards (RLVR) trained with group relative policy optimization (GRPO), the KL-free reward-advantage term st...
By Tommy Sha, Skylar Zhai, Siqi Zhao