arXiv:2607. 24953v1 Announce Type: cross Abstract: Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-bit oating point (FP4) remains challenging due to instability during optimization.
By Mehdi Rahimifar, Amin Darabi, Mehran Taghian Jazi, Xing Huang, Yao Wang, Zhijun Tu, Yufei Cui, Yunke Peng, Hongliang Li
arXiv:2603. 10444v2 Announce Type: replace-cross Abstract: FP4 training promises substantial memory and compute savings for large language models, but remains fragile because blockwise quantization is dictated by extreme activation magnitudes, which inflate dynamic range and compress long-tail signals.
By Hengjie Cao, Zhendong Huang, Mengyi Chen, Yifeng Yang, Fang Dong, Anrui Chen, Ruijun Huang, Xin Zhang, Mingzhi Dong, Yujiang Wang, Jinlong Hou, Qin Lv, Robert P. Dick, Yuan Cheng, Tun Lu, Fan Yang, Yixuan Chen, Li Shang
arXiv:2607. 26515v1 Announce Type: new Abstract: We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision.
By Hei Yi Mak, Shadan Golestan, Hoang Le, Mehran Taghian Jazi, Yunke Peng, Yaoyuan Wang, Yao Wang, Junsong Wang, Tianchi Hu, Fengchen He, Guipeng Hu, Tanzila Rahman, Anandharaju Durai Raju
The paper introduces Hardware‑Aware FP4 FlashAttention‑4, a method that leverages NVIDIA’s Blackwell 4‑bit floating‑point (FP4) tensor cores for attention mechanisms. It presents two key techniques: Direct‑P, which maps attention scores directly to FP4 probabilities for noncausal inference, achieving up to 2.13× the bfloat16 forward throughput on an NVIDIA GB200; and a causal path that reconstructs probabilities from quantized queries and keys while using 8‑bit floating‑point (FP8) gradients, accelerating a full single‑GPU 8‑billion‑parameter update by up to 1.14×. The authors also note that distributed training with matched FP8 probabilities and values diverges for every tested MXFP4 probability/value trajectory.
By Robert Hu
arXiv:2606. 15682v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong problem-solving through long chain-of-thought, but their deployment is constrained by the high cost of full-precision inference and growing KV cache footprints.
By Janghwan Lee, Sihwa Lee, Jinseok Kim, Yongjik Kim, Jieun Lim, Jinwook Oh, Jungwook Choi
arXiv:2508. 18224v3 Announce Type: replace-cross Abstract: Recent advances in sparse attention mechanisms have demonstrated strong potential for reducing the computational cost of long-context training and inference in large language models (LLMs).
By Ran Yan, Youhe Jiang, Zhuoming Chen, Haohui Mai, Beidi Chen, Binhang Yuan
The paper introduces Hardware‑Aware FP4 FlashAttention‑4, which optimizes attention mechanisms for NVIDIA’s Blackwell 4‑bit floating‑point (FP4) tensor cores. By employing Direct‑P for noncausal inference and a causal path that forwards quantized scores into the backward pass, the method achieves up to 2.13× the bfloat16 forward throughput on an NVIDIA GB200. The causal approach also reconstructs probabilities from saved quantized queries and keys, using 8‑bit floating‑point (FP8) gradients to accelerate a full single‑GPU 8‑billion‑parameter update by up to 1.14×, while distributed training with FP8 probabilities and values shows divergent trajectories compared to tested MXFP4 setups.
arXiv:2607. 04302v1 Announce Type: cross Abstract: We present HiFA4, a post-training operator-level design that executes both QK^T and PV in FlashAttention as 4-bit HIF4 Cube GEMMs for LLM inference on Ascend NPUs, while maintaining the online softmax state in FP16.
By Hui Dong, Yanzhao Li, Jie Gao, Chunlu Li, Zhiyuan Zhang, Yupeng Sun, Zhenyuan Chen, Zhiqiang Zou
arXiv:2609.38166v1 Announce Type: cross
Abstract: Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Atte...
By Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica
Block Sparse Flash Attention (BSFA) is a drop‑in replacement for FlashAttention that speeds up long‑context inference by pruning about 50% of computation and memory transfers. It selects the top‑k most important value blocks for each query using exact query‑key similarities and calibrated per‑layer, per‑head thresholds, requiring only a one‑time training‑free calibration. On Llama‑3.1‑8B, BSFA delivers up to 1.13× speedup on LongBench with a 1.1% accuracy drop and up to 1.24× on Needle‑in‑a‑Haystack retrieval with a 1% drop, while the attention kernel itself accelerates by up to 1.38×.
By Daniel Ohayon, Itay Lamprecht, Itay Hubara, Israel Cohen, Daniel Soudry, Noam Elata
arXiv:2601. 22813v2 Announce Type: replace Abstract: The NVFP4 lower-precision format, supported in hardware by NVIDIA Blackwell GPUs, promises to allow, for the first time, end-to-end fully-quantized pre-training of massive models such as LLMs.
By Andrei Panferov, Erik Schultheis, Soroush Tabesh, Dan Alistarh
arXiv:2609.39223v2 Announce Type: new
Abstract: Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive p...
By Weili Xu, Jisen Li, Yuqing Jian, Chenxi Li, Zhizhou Sha, Yifan Yu, Qingyang Wu, Chenfeng Xu, Zhongzhu Zhou, Tianyi Zhang, Ben Athiwaratkun