arXiv AI By Siddharth Aphale, Kelly Liu

SFT Overtraining Predicts Rank Inversion via Entropy Collapse Under RLVR

Read the original on arXiv AI →

arXiv:2606. 18487v1 Announce Type: cross Abstract: The standard heuristic of selecting the SFT checkpoint with the highest pass@1 for GRPO can fail when SFT compresses the rollout distribution.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 1

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

The paper investigates why reinforcement learning with verifiable rewards (RLVR) reduces the diversity of solutions in reasoning tasks. By analyzing the Countdown task, the authors show that RLVR contracts the solution space mainly at the entrance—before the first arithmetic operation—causing a 67% drop in solution coverage. They demonstrate that providing an unselected entrance prefix or applying entrance‑targeted interventions can restore or even improve coverage without harming accuracy.

By Qiancheng Zhou, Ruizhe Li
arXiv AI
Sep 2

Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO

The study investigates whether post‑training methods—GRPO, SFT, and DPO—improve language models’ ability to follow prompt evidence that conflicts with memorized knowledge. By comparing nine training variants across different scales and families, the authors find that grounding gains are modest for GRPO, moderate for Conflict‑SFT, and near‑ceiling for DPO, but all largely rely on the same causal attention‑head set present in the starting checkpoint. Removing the starting‑model grounding direction suppresses these gains, while adding it back recovers a significant portion of DPO’s improvement, indicating that existing model machinery drives most of the observed gains.

By Prakhar Gupta, Vaibhav Gupta
Hugging Face Trending Papers
Jul 30

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts

Reinforcement learning from verifiable rewards (RLVR) for mathematical reasoning suffers from a structural blind spot: on "cliff" prompts-those on which every sampled rollout in a group fails-the group-normalized advantage is identically zero, so GRPO produces no gradient on precisely the prompts at the frontier of the model's capability. We introduce LoRA Scaffolded Policy Optimization (LSPO), a sampling-time mechanism that recovers this lost gradient.