Benford's Law as a Distributional Prior for Post-Training Quantization of Large Language Models
Read the original on arXiv Machine Learning →The paper introduces BenQ, a data‑free post‑training quantization (PTQ) method that leverages Benford’s Law to guide the construction of a log‑spaced codebook for transformer weights. BenQ selectively applies this codebook to linear layers while preserving higher precision for LayerNorm parameters, achieving consistent 4‑bit group‑wise PTQ improvements over uniform RTN and competitive results with NF4 across various models and tasks. The authors also explore dynamic activation quantization, noting that log‑spaced grids can mitigate RTN failures but that handling outliers remains crucial for reliable low‑bit activation PTQ.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.