arXiv Machine Learning By Taiheng Pan

Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models

Read the original on arXiv Machine Learning →

The paper introduces PTH (Probe The Harness), a set of checks designed to expose hidden details in experimental setups that can alter the ranking of stale-data reinforcement learning methods for language models. By applying PTH to a comparison between SAN and truncated importance sampling (TIS), the authors demonstrate that subtle harness configurations—such as how PPO ratios are computed, data seeding, replay queue reuse, and loss normalisation—can reverse the observed performance order. The study provides a detailed signature of each influencing factor, reference results for TIS and uncorrected GRPO, and a checklist to ensure fair comparisons.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
2d ago

What Must Replay Preserve? Separating Correctable Bias from Class Correspondence

The paper investigates what information must be preserved in replay buffers for class‑incremental learning. By treating cached predictions as temporally heterogeneous supervision, the authors separate classes known at storage time from those learned later, and evaluate the impact of deleting logit matching. Experiments on CIFAR‑100 with DER++ show that a simple task‑level offset can largely correct the cost of removing later‑class matching, while the cost of disrupting class correspondence remains.

By BoRen Deng, Xiangyue Ma, Chenglong Li, Xiaoting Du
arXiv Machine Learning
Sep 18

Score Centering Stabilizes Off-policy Reinforcement Learning

The paper introduces a method called score centering to address the training‑inference mismatch (TIM) that destabilizes reinforcement learning for large language models. By adding an additive correction term that cancels drift between training and inference engines, score centering stabilizes RL and can match or surpass importance‑sampling techniques, especially as model size and mismatch severity increase. The approach also composes with importance sampling, yielding further performance gains in staleness experiments.

By Martin Marek, Max Ryabinin