The paper introduces PTH (Probe The Harness), a set of checks designed to expose hidden details in experimental setups that can alter the ranking of stale-data reinforcement learning methods for language models. By applying PTH to a comparison between SAN and truncated importance sampling (TIS), the authors demonstrate that subtle harness configurations—such as how PPO ratios are computed, data seeding, replay queue reuse, and loss normalisation—can reverse the observed performance order. The study provides a detailed signature of each influencing factor, reference results for TIS and uncorrected GRPO, and a checklist to ensure fair comparisons.
By Taiheng Pan
arXiv:2609.13922v1 Announce Type: new
Abstract: Minibatch persistency reuses data instead of reading it: rather than drawing a fresh minibatch at every optimizer step, it takes K consecutive steps on...
By Matteo Fischetti
arXiv:2609.11149v3 Announce Type: replace-cross
Abstract: How fast does a language model degrade when trained on its own outputs? Theory traces it to gradually accumulating errors, while experiments...
By Yangze Liu, Zhongyi Han
The paper introduces Temperon, a training strategy that uses plain SGD for the first 43% of the epoch budget and then hands off to a SAM‑wrapped Muon refiner for the remaining training. On datasets such as CIFAR‑10/100, SVHN, and Tiny ImageNet, Temperon achieves the same or better accuracy as full‑time SAM while reaching key performance targets faster and at lower cost. Ablation studies show that the Muon refiner contributes the majority of the performance gain, while the initial SGD explorer and its restarts add negligible benefit.
By Stamatis Mastromichalakis
arXiv:2608. 07911v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard.
By Yu Zhang
arXiv:2605. 17314v2 Announce Type: replace-cross Abstract: We consider whether off-policy experience from a smaller, weaker model can elicit capability in a stronger learner that on-policy RL fine-tuning (e.
By Wei Deng
arXiv:2608.07911v4 Announce Type: replace
Abstract: Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard. This makes expert c...
By Yu Zhang
arXiv:2610.09519v1 Announce Type: new
Abstract: Production machine-learning models are derived artifacts of time-bounded training snapshots: a deployed model is a materialized view over a training cu...
By Amit Rajula
The paper investigates what information must be preserved in replay buffers for class‑incremental learning. By treating cached predictions as temporally heterogeneous supervision, the authors separate classes known at storage time from those learned later, and evaluate the impact of deleting logit matching. Experiments on CIFAR‑100 with DER++ show that a simple task‑level offset can largely correct the cost of removing later‑class matching, while the cost of disrupting class correspondence remains.
By BoRen Deng, Xiangyue Ma, Chenglong Li, Xiaoting Du
arXiv:2604.11508v3 Announce Type: replace
Abstract: Fine-tuning a pretrained classifier leaves some samples reliably learned and others cycling between correct and incorrect. Curriculum learning, dat...
By Miit Daga, Swarna Priya Ramu
arXiv:2606. 27472v1 Announce Type: cross Abstract: Large language model (LLM) agents operate over long, multi-session interactions in which facts change: a user moves, a price updates, a plan is revised.
By Vedant Patel
The paper evaluates Certo, a small non‑generative decision model that scores candidate actions based on text. It compares a joint scorer that processes state, rules, and candidates together with a cacheable encoder that pre‑encodes candidates to reduce cost. Experiments show the cacheable approach loses rule sensitivity, while targeted counterfactual supervision can recover performance on synthetic tasks; however, on real rules the joint scorer still outperforms the cacheable version, and cross‑domain mixtures do not improve accuracy.
By Dushyant Rajput (AltSlate Labs LLP), Nirdesh Chauhan (AltSlate Labs LLP), Siddharth Kosaraju (AltSlate Labs LLP)