arXiv Machine Learning

RDQ: Residual Distribution Quantization for Large Language Models

arXiv:2607. 10137v1 Announce Type: new Abstract: Post-training quantization (PTQ) of large language models degrades sharply below 4-bit precision.

arXiv Machine Learning
Sep 22

PRQuant: Permutation Residual Quantization for Low-Overhead Inference

PRQuant introduces a training‑free, low‑overhead method for low‑bit quantization of linear layers by permuting input channels that cause the largest quantization error into contiguous tail blocks and precomputing residual weight sub‑tensors. The approach eliminates the need for online gathering during inference, converting scattered residual compensation into a regular tail‑augmented GEMM and thereby reducing latency. Experiments show that PRQuant lowers down‑projection reconstruction error and outperforms standard MXFP4 and other post‑training quantization baselines on five downstream benchmarks, improving accuracy by up to 1.24 points on Qwen3‑4B‑Instruct‑2507.

By Peiran Wang, Anqi Wang, Jiaying Zhao, Huiwen Yang, Zhenyu Ming, Rongqian Wang, Yiwu Yao, Kun Tian, Xin Yao, Gong Zhang, Fan Yang, Zhongyi Huang
arXiv Machine Learning
Sep 14

Benford's Law as a Distributional Prior for Post-Training Quantization of Large Language Models

The paper introduces BenQ, a data‑free post‑training quantization (PTQ) method that leverages Benford’s Law to guide the construction of a log‑spaced codebook for transformer weights. BenQ selectively applies this codebook to linear layers while preserving higher precision for LayerNorm parameters, achieving consistent 4‑bit group‑wise PTQ improvements over uniform RTN and competitive results with NF4 across various models and tasks. The authors also explore dynamic activation quantization, noting that log‑spaced grids can mitigate RTN failures but that handling outliers remains crucial for reliable low‑bit activation PTQ.

By Arthur Negr\~ao, Pedro Silva, Vander L. S. Freitas, Gladston Moreira, Eduardo Luz
arXiv Machine Learning
2d ago

QATFactory: A Versatile, Deployment-Aligned Framework for Quantization-aware Training and Distillation of LLMs

arXiv:2609.39223v2 Announce Type: new Abstract: Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive p...

By Weili Xu, Jisen Li, Yuqing Jian, Chenxi Li, Zhizhou Sha, Yifan Yu, Qingyang Wu, Chenfeng Xu, Zhongzhu Zhou, Tianyi Zhang, Ben Athiwaratkun