arXiv AI

An 84-Format Numeric Catalog with Bit-Exact Conformance Vectors: A Vendor-Neutral Reference for FP8, BF16, MXFP4, and Microscaling Formats

arXiv:2606. 09686v1 Announce Type: cross Abstract: Numeric format proliferation in machine learning hardware -- FP8 (E4M3 and E5M2), BF16, MXFP4, microscaling block formats, and dozens of research variants -- has outpaced the availability of vendor-neutral, bit-exact reference material.

arXiv AI
Sep 7

Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats

The paper introduces the Golden Ruler, a catalog of 109 numeric formats for machine learning hardware, including FP8, BF16, MXFP4, and microscaling block formats. It provides six bit‑exact conformance packs that map to IEEE P3109 v3.2.0 standards, each as a self‑contained JSON document with a SHA‑256 fingerprint and an anchor vector for cross‑pack sanity checks. The work cross‑validates these packs against ml_dtypes 0.5.4 and documents any divergences as spec‑permitted gaps, offering a vendor‑neutral reference for engineers.

By Dmitrii Vasilev
arXiv Machine Learning
Sep 11

Numbat: Building and Verifying a Self-Contained Machine-Learning Stack

The paper introduces Numbat, a self‑contained machine‑learning stack implemented entirely in Zig with no external runtime dependencies. It covers tensor computation, automatic differentiation, neural‑network modules, mixed precision, multi‑GPU training, data loading, and monitoring, and exposes a stable C ABI with over 1,400 entry points and bindings for six languages. The authors verify the stack against a reference implementation at multiple levels, uncovering ten silent recipe divergences, and demonstrate its practical capability by training a 25.9M‑parameter YOLOv8m detector on COCO 2017, achieving a competitive mAP score and matching single‑GPU performance.

By Thang Tran (CloudKites AI Lab, New South Wales, Australia), Lan Dang (Monash Business School, Monash University, Victoria, Australia)
arXiv Machine Learning
1d ago

Format-Aware Fusion for Fast FP4 Pretraining

The paper introduces format‑aware fusion, a method that co‑designs quantization producers with their scale domains and consumer layouts to fully exploit four‑bit floating‑point (FP4) Tensor Cores. Using this approach, the authors pretrain the Llama‑3‑family 8B model on 160 billion tokens, achieving up to 37.9 K tokens/s/GPU—significantly higher than standard bfloat16 or Transformer Engine FP4 baselines. The study demonstrates that FP4 performance depends on the interplay of scaling, operand packing, layout, and execution path, with downstream task rankings diverging from training‑loss rankings.

By Robert Hu