arXiv AI By Xiaocan Li, Shiliang Wu, Zheng Shen

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor

Read the original on arXiv AI →

arXiv:2605. 20402v3 Announce Type: replace-cross Abstract: MXFP4 arithmetic can dramatically accelerate reinforcement learning (RL) post-training of large language models (LLMs), yet the quantization error introduces severe accuracy degradation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 20

QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides

arXiv:2607. 15810v1 Announce Type: new Abstract: Rollout generation is a major bottleneck in Reinforcement Learning (RL) for Mixture-of-Experts (MoE) Large Language Models, motivating low-precision rollout acceleration such as FP8.

By Zhengyang Zhuge, Hao Yu, Xin Wang, Zheng Li, Yizhong Cao, Dayiheng Liu, Jianwei Zhang
arXiv Machine Learning
Sep 22

Towards Full Pipeline FP8 Reinforcement Learning for LLMs

The paper introduces Calibrated Clipping, a dynamic method to align FP8 quantization bounds with high‑precision BF16 distributions, thereby mitigating training instability in full‑pipeline FP8 reinforcement learning for large language models. It identifies that compounded FP8 noise distorts importance ratios, causing entropy surges and garbled outputs. Experiments across GRPO and DAPO algorithms on 8B‑32B models show the technique restores performance to BF16 levels.

By Fanchao Chen, Ziheng Jiang, Ziyun Wei, Zheng Zhong, Du Li, Chi Zhang, Haibin Lin, Shivaram Venkataraman
arXiv Machine Learning
Jul 30

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

arXiv:2607. 26515v1 Announce Type: new Abstract: We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision.

By Hei Yi Mak, Shadan Golestan, Hoang Le, Mehran Taghian Jazi, Yunke Peng, Yaoyuan Wang, Yao Wang, Junsong Wang, Tianchi Hu, Fengchen He, Guipeng Hu, Tanzila Rahman, Anandharaju Durai Raju
arXiv AI
Sep 2

REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

REAL-Q introduces a new post‑training quantization approach for large language models that replaces the traditional single closed‑form second‑order solver with a fine‑grained, dynamic block‑wise gradient descent applied after every 128‑column block. By aligning the surrogate loss with the end‑to‑end objective and using a sliding window for smooth cross‑layer transitions, REAL‑Q mitigates error propagation and information misalignment. Experiments on LLaMA‑3.1 and Qwen3 show up to ~49% reduction in end‑to‑end KL divergence compared to state‑of‑the‑art methods.

By Qian Zhang, Yaoming Li, Zhewen Tan, Yanshu Wang, Heng Lu, Kun Su, Zongwei Lv, Wenhan Yu, Yongge Ma, Yinjun Han, Ruikuang Liu, Tong Yang