arXiv AI

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning

arXiv:2607. 24062v1 Announce Type: cross Abstract: Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between training and inference.

arXiv Machine Learning
Sep 22

Towards Full Pipeline FP8 Reinforcement Learning for LLMs

The paper introduces Calibrated Clipping, a dynamic method to align FP8 quantization bounds with high‑precision BF16 distributions, thereby mitigating training instability in full‑pipeline FP8 reinforcement learning for large language models. It identifies that compounded FP8 noise distorts importance ratios, causing entropy surges and garbled outputs. Experiments across GRPO and DAPO algorithms on 8B‑32B models show the technique restores performance to BF16 levels.

By Fanchao Chen, Ziheng Jiang, Ziyun Wei, Zheng Zhong, Du Li, Chi Zhang, Haibin Lin, Shivaram Venkataraman
arXiv Machine Learning
Jun 30

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

arXiv:2606. 29526v1 Announce Type: new Abstract: Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse.

By Jing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, Ju Huang, Wenbo Su, Jinyi Liu, Yan Zheng, Jianye Hao, Bo Zheng
arXiv Machine Learning
Sep 18

Score Centering Stabilizes Off-policy Reinforcement Learning

The paper introduces a method called score centering to address the training‑inference mismatch (TIM) that destabilizes reinforcement learning for large language models. By adding an additive correction term that cancels drift between training and inference engines, score centering stabilizes RL and can match or surpass importance‑sampling techniques, especially as model size and mismatch severity increase. The approach also composes with importance sampling, yielding further performance gains in staleness experiments.

By Martin Marek, Max Ryabinin
arXiv AI
Sep 10

Evaluating the Scalability and Adversarial Generalization of GRPO-Trained NLI Models

The paper evaluates the scalability and adversarial generalization of Natural Language Inference (NLI) models trained with Group Relative Policy Optimization (GRPO) for Chain-of-Thought learning. By fine‑tuning 7B, 14B, and 32B language models with LoRA and QLoRA, the authors show strong performance on standard and adversarial NLI benchmarks, with the 32B model outperforming supervised baselines on adversarial sets. Using AWQ quantization, the 32B model fits within 22 GB of CUDA memory, demonstrating a scalable, practical framework for robust NLI without sacrificing inference quality.

By Pablo Miralles-Gonz\'alez, Javier Huertas-Tato, Alejandro Mart\'in, David Camacho
arXiv Machine Learning
Jun 8

Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning

arXiv:2606. 06673v1 Announce Type: new Abstract: Sparse rewards and heterogeneous task sequences remain persistent challenges in Reinforcement Learning (RL), often resulting in slow convergence, weak generalization, and inefficient exploration.

By Ujjwal Bhatta, Utsabi Dangol, Sumaly Bajracharya, Rodrigue Rizk, KC Santosh