arXiv Machine Learning

Format-Aware Fusion for Fast FP4 Pretraining

The paper introduces format‑aware fusion, a method that co‑designs quantization producers with their scale domains and consumer layouts to fully exploit four‑bit floating‑point (FP4) Tensor Cores. Using this approach, the authors pretrain the Llama‑3‑family 8B model on 160 billion tokens, achieving up to 37.9 K tokens/s/GPU—significantly higher than standard bfloat16 or Transformer Engine FP4 baselines. The study demonstrates that FP4 performance depends on the interplay of scaling, operand packing, layout, and execution path, with downstream task rankings diverging from training‑loss rankings.

arXiv Machine Learning
Sep 3

UE5M3 FP4 Block Scaling for Stable Language Model Pretraining

The paper introduces a new 4‑bit floating‑point (FP4) pretraining approach that pairs E2M1 payloads with unsigned E5M3 block scales, enabling periodic tensor scaling and selective stochastic rounding while eliminating the randomized Hadamard transform. Using this method, the authors pretrained a Nemotron‑H 8B model on nearly 190 billion tokens, achieving lower training and validation losses compared to NVIDIA’s Transformer Engine. The approach also improves inference performance and demonstrates a 21.2 % increase in token throughput when certain optimizations are removed.

By Robert Hu, Carlo Luschi, Paul Balanca
arXiv Machine Learning
Jun 15

Realizing Native INT8 Compute for Diffusion Transformers on Consumer GPUs: A Fused INT8 GEMM Kernel for Ideogram 4.0

arXiv:2606. 14598v1 Announce Type: new Abstract: Post-training INT8 (W8A8) quantization of diffusion transformers is widely deployed as a speed optimization, yet on consumer Ampere GPUs it is frequently slower than the FP8 and NF4 alternatives it is meant to beat.

By Ali Asaria, Tony Salomone, Deep Gandhi
arXiv AI
Jun 19

Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

arXiv:2606. 20381v1 Announce Type: new Abstract: FP4 training promises substantial reductions in memory and computation cost for LLM pretraining, yet current FP4 hardware paths and recipes, including NVIDIA Blackwell/Rubin-class systems and AMD MI350-series GPUs, remain centered on E2M1 data elements.

By Qian Zhao, Kunlong Chen, Changxin Tian, Zhonghui Jiang, Haitao Zhang, Chaofan Yu, Peijie Jiang, Mingliang Gong, Jia Liu, Ziqi Liu, Zhiqiang Zhang, Jun Zhou
arXiv Machine Learning
2d ago

QATFactory: A Versatile, Deployment-Aligned Framework for Quantization-aware Training and Distillation of LLMs

arXiv:2609.39223v2 Announce Type: new Abstract: Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive p...

By Weili Xu, Jisen Li, Yuqing Jian, Chenxi Li, Zhizhou Sha, Yifan Yu, Qingyang Wu, Chenfeng Xu, Zhongzhu Zhou, Tianyi Zhang, Ben Athiwaratkun
Hugging Face Trending Papers
Sep 3

Hardware-Aware FP4 FlashAttention-4

The paper introduces Hardware‑Aware FP4 FlashAttention‑4, which optimizes attention mechanisms for NVIDIA’s Blackwell 4‑bit floating‑point (FP4) tensor cores. By employing Direct‑P for noncausal inference and a causal path that forwards quantized scores into the backward pass, the method achieves up to 2.13× the bfloat16 forward throughput on an NVIDIA GB200. The causal approach also reconstructs probabilities from saved quantized queries and keys, using 8‑bit floating‑point (FP8) gradients to accelerate a full single‑GPU 8‑billion‑parameter update by up to 1.14×, while distributed training with FP8 probabilities and values shows divergent trajectories compared to tested MXFP4 setups.

arXiv Machine Learning
Sep 4

Hardware-Aware FP4 FlashAttention-4

The paper introduces Hardware‑Aware FP4 FlashAttention‑4, a method that leverages NVIDIA’s Blackwell 4‑bit floating‑point (FP4) tensor cores for attention mechanisms. It presents two key techniques: Direct‑P, which maps attention scores directly to FP4 probabilities for noncausal inference, achieving up to 2.13× the bfloat16 forward throughput on an NVIDIA GB200; and a causal path that reconstructs probabilities from quantized queries and keys while using 8‑bit floating‑point (FP8) gradients, accelerating a full single‑GPU 8‑billion‑parameter update by up to 1.14×. The authors also note that distributed training with matched FP8 probabilities and values diverges for every tested MXFP4 probability/value trajectory.

By Robert Hu
arXiv Machine Learning
Sep 23

Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

The paper addresses the problem of non‑deterministic outputs from large language models (LLMs) when run on different GPU architectures, caused by floating‑point non‑associativity and hardware‑dependent kernel choices. It proposes a set of fixed‑configuration fused‑upcast GEMM kernels that load 16‑bit weights, upcast to FP32, and perform IEEE‑754 compliant reductions in a problem‑shape‑dependent order, ensuring identical linear‑layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs. The new approach achieves 1.17–3.1× faster end‑to‑end performance than existing solutions and halves weight‑memory traffic while maintaining cross‑architecture reproducibility.

By Liam Cooper, Shinnung Jeong, Hyeran Jeon, Jeffrey Young, Hyesoon Kim