arXiv:2609.36953v1 Announce Type: new
Abstract: Production RL for language models lets the sampler fall behind the learner and repairs the resulting mismatch with a truncated importance weight. We as...
By Taiheng Pan
arXiv:2609.06107v1 Announce Type: new
Abstract: Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which do...
By Hao Liang, Mingrui Chen, Hengyi Feng, Meiyi Qiang, Wentao Zhang
The paper investigates what information must be preserved in replay buffers for class‑incremental learning. By treating cached predictions as temporally heterogeneous supervision, the authors separate classes known at storage time from those learned later, and evaluate the impact of deleting logit matching. Experiments on CIFAR‑100 with DER++ show that a simple task‑level offset can largely correct the cost of removing later‑class matching, while the cost of disrupting class correspondence remains.
By BoRen Deng, Xiangyue Ma, Chenglong Li, Xiaoting Du
arXiv:2607. 10848v1 Announce Type: new Abstract: Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize off-policy updates.
By Xiangxin Zhou, Jiarui Yao, Penghui Qi, Bowen Ping, Jiaqi Tang, Haonan Wang, Tianyu Pang
The paper introduces a method called score centering to address the training‑inference mismatch (TIM) that destabilizes reinforcement learning for large language models. By adding an additive correction term that cancels drift between training and inference engines, score centering stabilizes RL and can match or surpass importance‑sampling techniques, especially as model size and mismatch severity increase. The approach also composes with importance sampling, yielding further performance gains in staleness experiments.
By Martin Marek, Max Ryabinin
arXiv:2606. 15333v1 Announce Type: cross Abstract: LLM unlearning has emerged as a cost-effective alternative to full retraining for removing hazardous knowledge from pretrained models while preserving general utility.
By Zirui Pang, Chenlong Zhang, Haosheng Tan, Zhuoran Jin, Jiaheng Wei, Zixin Zhong
arXiv:2610.00385v1 Announce Type: cross
Abstract: Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and d...
By Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Tianshu Fu, Daren Zha, Jun Xiao
The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.
By Ayoub Kirouane, Christos Petrocheilos
arXiv:2608.23830v1 Announce Type: cross
Abstract: RL has emerged as a powerful paradigm for enhancing the instruction following capabilities of LLMs. While existing training recipes achieve substanti...
By Mian Zhang, Yueqin Yin, Kaiyu He, Peilin Wu, Xinlu Zhang, Mingyuan Zhou, Zhiyu Zoey Chen
The paper introduces COPC, a Coupled Off-Policy Correction method for asynchronous reinforcement learning of large language models. COPC coordinates policy-side and advantage-side corrections by combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals, addressing both policy mismatch and advantage staleness. Experiments show COPC outperforms existing asynchronous baselines on tool-integrated mathematical reasoning and search tasks, while maintaining training stability and minimal overhead.
By Zicheng Hu, Zhijian Zhou, Xuan Zhang, Yuchen Liu, Cheng Chen, Yuan Li, Qi Gu, Yan Feng, Hongyan Hao, Chao Qu
UpgradeBench is a decision‑centric longitudinal benchmark that evaluates how fine‑tuned language‑model specialists should be handled when new base‑model releases occur. It covers four consecutive Qwen releases, a continuation checkpoint, six tasks, two model sizes, and OLMo checkpoints with known training lineage, and examines whether retraining, adapter transfer, or other recovery strategies improve specialist performance. The benchmark reveals that upgrade gains vary by task and release interval, that direct adapter copying is sensitive to pretraining distance, and that teacher relabeling can recover specialists without new annotations.
"whyItMatters":"The study provides actionable insights into the cost‑effective management of specialist models across model releases, showing how to balance retraining effort with performance gains."
By Ye Chen, Weining Zhang
arXiv:2605. 02909v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large language models (LLMs).
By Kazuki Egashira, Mark Vero, Jasper Dekoninck, Florian E. Dorner, Robin Staab, Martin Vechev