Cool the Sampler, Not the Learner: Sampling Temperature Moves the Staleness Cliff of Importance-Corrected GRPO
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper introduces PTH (Probe The Harness), a set of checks designed to expose hidden details in experimental setups that can alter the ranking of stale-data reinforcement learning methods for language models. By applying PTH to a comparison between SAN and truncated importance sampling (TIS), the authors demonstrate that subtle harness configurations—such as how PPO ratios are computed, data seeding, replay queue reuse, and loss normalisation—can reverse the observed performance order. The study provides a detailed signature of each influencing factor, reference results for TIS and uncorrected GRPO, and a checklist to ensure fair comparisons.
arXiv:2609.13922v1 Announce Type: new Abstract: Minibatch persistency reuses data instead of reading it: rather than drawing a fresh minibatch at every optimizer step, it takes K consecutive steps on...
arXiv:2609.11149v3 Announce Type: replace-cross Abstract: How fast does a language model degrade when trained on its own outputs? Theory traces it to gradually accumulating errors, while experiments...
The paper introduces Temperon, a training strategy that uses plain SGD for the first 43% of the epoch budget and then hands off to a SAM‑wrapped Muon refiner for the remaining training. On datasets such as CIFAR‑10/100, SVHN, and Tiny ImageNet, Temperon achieves the same or better accuracy as full‑time SAM while reaching key performance targets faster and at lower cost. Ablation studies show that the Muon refiner contributes the majority of the performance gain, while the initial SGD explorer and its restarts add negligible benefit.
arXiv:2608. 07911v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard.
arXiv:2605. 17314v2 Announce Type: replace-cross Abstract: We consider whether off-policy experience from a smaller, weaker model can elicit capability in a stronger learner that on-policy RL fine-tuning (e.