arXiv Machine Learning By Zhanming Zhang, Vinoth Selvendran

Sharpen Before You Adapt: Data-Free Entry-State Sharpening for Test-Time Reinforcement Learning

Read the original on arXiv Machine Learning →

The paper introduces entry-state sharpening, a data‑free pre‑training step that prepares a language model’s checkpoint in a sharper, lower‑entropy state before test‑time reinforcement learning (TTRL). By reducing policy entropy, the model can more efficiently use its limited adaptation budget, leading to higher endpoint conversion efficiency across tasks such as MATH, GPQA, and AMC. Experiments with different data‑free objectives (e.g., R‑Zero vs. SPIRAL) demonstrate that the choice of pre‑training objective strongly influences the checkpoint’s readiness for TTRL, and a label‑free self‑distillation intervention can further sharpen the entry state.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

Learning from the Near Future: Temporal Self-Distillation for RLVR

The paper introduces temporal self‑distillation for reinforcement learning with verifiable rewards (RLVR), proposing that a policy can learn from a stronger future checkpoint of itself. Two methods—Near‑Future Policy Optimization (NPO) and Near‑Future Policy Distillation (NPD)—use verified future‑self trajectories and token‑level transfer, respectively, while AutoNPO adaptively selects the optimal future checkpoint. Experiments on eight image‑text benchmarks show that near‑future teachers yield higher performance than far‑future ones, indicating that the balance between new capability and learner compatibility is key.

By Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang
arXiv Machine Learning
Sep 1

The Intervention Gap in Latent World Models

The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.

By Donna Vakalis
arXiv AI
Sep 1

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

The paper investigates why reinforcement learning with verifiable rewards (RLVR) reduces the diversity of solutions in reasoning tasks. By analyzing the Countdown task, the authors show that RLVR contracts the solution space mainly at the entrance—before the first arithmetic operation—causing a 67% drop in solution coverage. They demonstrate that providing an unselected entrance prefix or applying entrance‑targeted interventions can restore or even improve coverage without harming accuracy.

By Qiancheng Zhou, Ruizhe Li