arXiv Machine Learning

Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models

The paper introduces PTH (Probe The Harness), a set of checks designed to expose hidden details in experimental setups that can alter the ranking of stale-data reinforcement learning methods for language models. By applying PTH to a comparison between SAN and truncated importance sampling (TIS), the authors demonstrate that subtle harness configurations—such as how PPO ratios are computed, data seeding, replay queue reuse, and loss normalisation—can reverse the observed performance order. The study provides a detailed signature of each influencing factor, reference results for TIS and uncorrected GRPO, and a checklist to ensure fair comparisons.

arXiv Machine Learning
2d ago

What Must Replay Preserve? Separating Correctable Bias from Class Correspondence

The paper investigates what information must be preserved in replay buffers for class‑incremental learning. By treating cached predictions as temporally heterogeneous supervision, the authors separate classes known at storage time from those learned later, and evaluate the impact of deleting logit matching. Experiments on CIFAR‑100 with DER++ show that a simple task‑level offset can largely correct the cost of removing later‑class matching, while the cost of disrupting class correspondence remains.

By BoRen Deng, Xiangyue Ma, Chenglong Li, Xiaoting Du
arXiv Machine Learning
Sep 18

Score Centering Stabilizes Off-policy Reinforcement Learning

The paper introduces a method called score centering to address the training‑inference mismatch (TIM) that destabilizes reinforcement learning for large language models. By adding an additive correction term that cancels drift between training and inference engines, score centering stabilizes RL and can match or surpass importance‑sampling techniques, especially as model size and mismatch severity increase. The approach also composes with importance sampling, yielding further performance gains in staleness experiments.

By Martin Marek, Max Ryabinin
arXiv Machine Learning
Aug 19

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.

By Ayoub Kirouane, Christos Petrocheilos
arXiv Machine Learning
1d ago

COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning

The paper introduces COPC, a Coupled Off-Policy Correction method for asynchronous reinforcement learning of large language models. COPC coordinates policy-side and advantage-side corrections by combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals, addressing both policy mismatch and advantage staleness. Experiments show COPC outperforms existing asynchronous baselines on tool-integrated mathematical reasoning and search tasks, while maintaining training stability and minimal overhead.

By Zicheng Hu, Zhijian Zhou, Xuan Zhang, Yuchen Liu, Cheng Chen, Yuan Li, Qi Gu, Yan Feng, Hongyan Hao, Chao Qu
arXiv AI
Aug 24

UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

UpgradeBench is a decision‑centric longitudinal benchmark that evaluates how fine‑tuned language‑model specialists should be handled when new base‑model releases occur. It covers four consecutive Qwen releases, a continuation checkpoint, six tasks, two model sizes, and OLMo checkpoints with known training lineage, and examines whether retraining, adapter transfer, or other recovery strategies improve specialist performance. The benchmark reveals that upgrade gains vary by task and release interval, that direct adapter copying is sensitive to pretraining distance, and that teacher relabeling can recover specialists without new annotations. "whyItMatters":"The study provides actionable insights into the cost‑effective management of specialist models across model releases, showing how to balance retraining effort with performance gains."

By Ye Chen, Weining Zhang