arXiv Machine Learning By Arthur Negr\~ao, Pedro Silva, Vander L. S. Freitas, Gladston Moreira, Eduardo Luz

Benford's Law as a Distributional Prior for Post-Training Quantization of Large Language Models

Read the original on arXiv Machine Learning →

The paper introduces BenQ, a data‑free post‑training quantization (PTQ) method that leverages Benford’s Law to guide the construction of a log‑spaced codebook for transformer weights. BenQ selectively applies this codebook to linear layers while preserving higher precision for LayerNorm parameters, achieving consistent 4‑bit group‑wise PTQ improvements over uniform RTN and competitive results with NF4 across various models and tasks. The authors also explore dynamic activation quantization, noting that log‑spaced grids can mitigate RTN failures but that handling outliers remains crucial for reliable low‑bit activation PTQ.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 29

Stable FP4 Training via Transposition-Invariant Block Quantization

arXiv:2607. 24953v1 Announce Type: cross Abstract: Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-bit oating point (FP4) remains challenging due to instability during optimization.

By Mehdi Rahimifar, Amin Darabi, Mehran Taghian Jazi, Xing Huang, Yao Wang, Zhijun Tu, Yufei Cui, Yunke Peng, Hongliang Li