arXiv Machine Learning
Sep 22

PRQuant: Permutation Residual Quantization for Low-Overhead Inference

PRQuant introduces a training‑free, low‑overhead method for low‑bit quantization of linear layers by permuting input channels that cause the largest quantization error into contiguous tail blocks and precomputing residual weight sub‑tensors. The approach eliminates the need for online gathering during inference, converting scattered residual compensation into a regular tail‑augmented GEMM and thereby reducing latency. Experiments show that PRQuant lowers down‑projection reconstruction error and outperforms standard MXFP4 and other post‑training quantization baselines on five downstream benchmarks, improving accuracy by up to 1.24 points on Qwen3‑4B‑Instruct‑2507.

By Peiran Wang, Anqi Wang, Jiaying Zhao, Huiwen Yang, Zhenyu Ming, Rongqian Wang, Yiwu Yao, Kun Tian, Xin Yao, Gong Zhang, Fan Yang, Zhongyi Huang