arXiv Machine Learning By Prateek Singh

RDQ: Residual Distribution Quantization for Large Language Models

Read the original on arXiv Machine Learning →

arXiv:2607. 10137v1 Announce Type: new Abstract: Post-training quantization (PTQ) of large language models degrades sharply below 4-bit precision.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 22

PRQuant: Permutation Residual Quantization for Low-Overhead Inference

PRQuant introduces a training‑free, low‑overhead method for low‑bit quantization of linear layers by permuting input channels that cause the largest quantization error into contiguous tail blocks and precomputing residual weight sub‑tensors. The approach eliminates the need for online gathering during inference, converting scattered residual compensation into a regular tail‑augmented GEMM and thereby reducing latency. Experiments show that PRQuant lowers down‑projection reconstruction error and outperforms standard MXFP4 and other post‑training quantization baselines on five downstream benchmarks, improving accuracy by up to 1.24 points on Qwen3‑4B‑Instruct‑2507.

By Peiran Wang, Anqi Wang, Jiaying Zhao, Huiwen Yang, Zhenyu Ming, Rongqian Wang, Yiwu Yao, Kun Tian, Xin Yao, Gong Zhang, Fan Yang, Zhongyi Huang