Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
arXiv:2505. 01043v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved impressive performance across various domains.
arXiv:2606. 31308v1 Announce Type: new Abstract: This paper investigates the capability of Large Language Models (LLMs) to detect and classify floating-point errors statically in software code.
arXiv:2505. 01043v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved impressive performance across various domains.
arXiv:2509. 26476v3 Announce Type: replace-cross Abstract: We study code-to-metric regression: predicting numeric outcomes of code executions, a challenging task due to the open-ended nature of programming languages.
arXiv:2606. 20128v1 Announce Type: cross Abstract: Benchmarks for LLM-generated GPU kernels (KernelBench, TritonBench, GEAK) score correctness through fixed-shape, small-sample allclose-style checks.
arXiv:2605. 10886v3 Announce Type: replace-cross Abstract: Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8.
arXiv:2603. 15510v2 Announce Type: replace Abstract: The synthesis of inductive loop invariants remains a critical bottleneck in automated program verification.
arXiv:2608. 12004v1 Announce Type: cross Abstract: In modern AI frameworks, GPU kernels are key to overall system performance.
The paper introduces Numbat, a self‑contained machine‑learning stack implemented entirely in Zig with no external runtime dependencies. It covers tensor computation, automatic differentiation, neural‑network modules, mixed precision, multi‑GPU training, data loading, and monitoring, and exposes a stable C ABI with over 1,400 entry points and bindings for six languages. The authors verify the stack against a reference implementation at multiple levels, uncovering ten silent recipe divergences, and demonstrate its practical capability by training a 25.9M‑parameter YOLOv8m detector on COCO 2017, achieving a competitive mAP score and matching single‑GPU performance.
arXiv:2602. 01027v2 Announce Type: replace Abstract: Mixed-precision quantization is a promising approach for compressing large language models under tight memory budgets.
The paper "LLM Inference in a Flash!" proposes an integer‑only quantization scheme and a dictionary‑based KV cache compression technique to enable large language model inference on compute‑in‑flash (CIF) devices. By eliminating floating‑point operations and reducing KV cache traffic through sparse dictionary coding, the authors achieve minimal accuracy loss while cutting dynamic KV cache traffic by 15× on Llama‑3.1‑8B and Qwen‑2.5‑7B models.
arXiv:2602. 10431v4 Announce Type: replace Abstract: Large language models (LLMs) demand substantial computational and memory resources, posing challenges for efficient deployment.
arXiv:2609.10226v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in...
The paper demonstrates that greedy decoding from large language models is not precision‑invariant: the same model, prompt, and decoding algorithm can produce different outputs when run in BF16 versus FP16 on identical hardware. Across six models (1.1B–7B parameters, four families, and 12B) and three benchmarks, 49–100 % of prompts diverge, with a single token flip often cascading into trajectory‑level divergence. The authors develop an empirical error‑propagation analysis that identifies the top‑two logit margin at the LM head as the key factor, and they propose a low‑overhead intervention—selective FP32 LM head recomputation—that improves exact agreement by 22–36 percentage points with less than 4 % latency overhead. "whyItMatters":"The findings reveal that precision choices can fundamentally alter model outputs, challenging the assumption of deterministic greedy decoding and highlighting the need for precision‑aware inference strategies."