Hugging Face Trending Papers

StagQ: Constraint-Driven Multi-Precision Weight Quantization for LLMs

Read the original on Hugging Face Trending Papers →

StagQ is a multi‑precision weight format for large language models that uses a 2‑bit group‑wise affine base followed by optional 1‑bit refinement planes. Each supported precision can be read as a prefix of the main stream, decoded via a shared affine map without per‑weight lookups, and a sparse side record stores the few weights that the grid handles poorly. Experiments show that StagQ outperforms baseline multi‑precision schemes on Llama‑3.1‑8B, Phi‑4, and OLMo‑2‑7B across various bit‑widths, and its GPU kernel is faster than baseline kernels for most shape‑precision combinations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Sep 23

Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

The paper addresses the problem of non‑deterministic outputs from large language models (LLMs) when run on different GPU architectures, caused by floating‑point non‑associativity and hardware‑dependent kernel choices. It proposes a set of fixed‑configuration fused‑upcast GEMM kernels that load 16‑bit weights, upcast to FP32, and perform IEEE‑754 compliant reductions in a problem‑shape‑dependent order, ensuring identical linear‑layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs. The new approach achieves 1.17–3.1× faster end‑to‑end performance than existing solutions and halves weight‑memory traffic while maintaining cross‑architecture reproducibility.

By Liam Cooper, Shinnung Jeong, Hyeran Jeon, Jeffrey Young, Hyesoon Kim