arXiv:2607. 15810v1 Announce Type: new Abstract: Rollout generation is a major bottleneck in Reinforcement Learning (RL) for Mixture-of-Experts (MoE) Large Language Models, motivating low-precision rollout acceleration such as FP8.
By Zhengyang Zhuge, Hao Yu, Xin Wang, Zheng Li, Yizhong Cao, Dayiheng Liu, Jianwei Zhang
The paper introduces Calibrated Clipping, a dynamic method to align FP8 quantization bounds with high‑precision BF16 distributions, thereby mitigating training instability in full‑pipeline FP8 reinforcement learning for large language models. It identifies that compounded FP8 noise distorts importance ratios, causing entropy surges and garbled outputs. Experiments across GRPO and DAPO algorithms on 8B‑32B models show the technique restores performance to BF16 levels.
By Fanchao Chen, Ziheng Jiang, Ziyun Wei, Zheng Zhong, Du Li, Chi Zhang, Haibin Lin, Shivaram Venkataraman
arXiv:2607. 26515v1 Announce Type: new Abstract: We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision.
By Hei Yi Mak, Shadan Golestan, Hoang Le, Mehran Taghian Jazi, Yunke Peng, Yaoyuan Wang, Yao Wang, Junsong Wang, Tianchi Hu, Fengchen He, Guipeng Hu, Tanzila Rahman, Anandharaju Durai Raju
arXiv:2608.20873v1 Announce Type: new
Abstract: Every way of teaching a deployed language model something new -- full fine-tuning, adapter merging, model editing -- replaces the released checkpoint,...
By Zifeng Liu, Zhiyong Du, Yaxin Lu, Yiming Mao, Zhenhe Wang, Wenqi Shi, Zhengkun Jing
arXiv:2607. 24062v1 Announce Type: cross Abstract: Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between training and inference.
By Wenwu Fan, Qihong Lin, Zhijie Xia, Zhuo Zheng, Sihao Wang, Qiang Chen, Liangsheng Zhu
REAL-Q introduces a new post‑training quantization approach for large language models that replaces the traditional single closed‑form second‑order solver with a fine‑grained, dynamic block‑wise gradient descent applied after every 128‑column block. By aligning the surrogate loss with the end‑to‑end objective and using a sliding window for smooth cross‑layer transitions, REAL‑Q mitigates error propagation and information misalignment. Experiments on LLaMA‑3.1 and Qwen3 show up to ~49% reduction in end‑to‑end KL divergence compared to state‑of‑the‑art methods.
By Qian Zhang, Yaoming Li, Zhewen Tan, Yanshu Wang, Heng Lu, Kun Su, Zongwei Lv, Wenhan Yu, Yongge Ma, Yinjun Han, Ruikuang Liu, Tong Yang