arXiv:2510. 04212v4 Announce Type: replace-cross Abstract: The pursuit of computational efficiency has driven the adoption of low-precision formats for training transformer models.
By Haiquan Qiu, Quanming Yao
The paper introduces a new 4‑bit floating‑point (FP4) pretraining approach that pairs E2M1 payloads with unsigned E5M3 block scales, enabling periodic tensor scaling and selective stochastic rounding while eliminating the randomized Hadamard transform. Using this method, the authors pretrained a Nemotron‑H 8B model on nearly 190 billion tokens, achieving lower training and validation losses compared to NVIDIA’s Transformer Engine. The approach also improves inference performance and demonstrates a 21.2 % increase in token throughput when certain optimizations are removed.
By Robert Hu, Carlo Luschi, Paul Balanca
arXiv:2608. 06177v1 Announce Type: new Abstract: Binary neural networks are very attractive for constrained deployment, enabling small footprint and low-power inference.
By Quentin Luquet de Saint-Germain, Massil Ait Abdeslam, Jean Pierre David
arXiv:2610. 01889v1 Announce Type: new Abstract: Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)?
By Yohan Chatelain (Krembil Centre for Neuroinformatics, CAMH, Toronto, Canada), Pablo de Oliveira Castro (Universite Paris-Saclay, UVSQ, LI-PaRAD, Versailles, France)
Stable 4-bit floating-point (FP4) pretraining is difficult because the E2M1 payload represents only a narrow range of magnitudes. NVIDIA's Transformer Engine \nv{} recipe addresses this with current-tensor scaling, a randomized Hadamard transform (RHT), and bfloat16 (BF16) final layers, adding work outside the FP4 matrix multiplications.
arXiv:2607. 23777v1 Announce Type: cross Abstract: The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data.
By Anuj Apte