← Back to all news
arXiv Machine Learning July 14, 2026 By Prateek Singh

RDQ: Residual Distribution Quantization for Large Language Models

Read the original on arXiv Machine Learning →

arXiv:2607. 10137v1 Announce Type: new Abstract: Post-training quantization (PTQ) of large language models degrades sharply below 4-bit precision.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

  • llms
  • efficiency
  • benchmarks

Related stories

Hugging Face Trending Papers
Jun 11

TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization

Large language models (LLMs) exhibit exceptional general language processing capabilities, but their memory and compute costs hinder deployment. Ternarization has emerged as a promising compression technique, offering significant reductions in model size and inference complexity.

llmsefficiency
More like this →
arXiv AI
Jul 7

ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs

arXiv:2601. 07475v2 Announce Type: replace-cross Abstract: The emergence of fine-grained numerical formats like NVFP4 presents new opportunities for efficient Large Language Model (LLM) inference.

By Haoqian Meng, Yilun Luo, Yafei Zhao, Wenyuan Liu, Peng Zhang, Xindian Ma
llmsefficiencybenchmarks
More like this →
arXiv AI
Jun 12

TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization

arXiv:2606. 13054v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit exceptional general language processing capabilities, but their memory and compute costs hinder deployment.

By Zhixiong Zhao, Zukang Xu, Zhixuan Chen, Xing Hu, Zhe Jiang, Dawei Yang
llmsefficiency
More like this →
arXiv Machine Learning
Jul 10

KronQ: LLM Quantization via Kronecker-Factored Hessian

arXiv:2607. 07964v1 Announce Type: new Abstract: Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining.

By Donghyun Lee, Yuhang Li, Ruokai Yin, Priyadarshini Panda
llmsefficiency
More like this →
arXiv AI
Jul 28

MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models

arXiv:2607. 23047v1 Announce Type: cross Abstract: Mixed-precision quantization improves the accuracy of post-training quantization by allocating higher bitwidths to sensitive layers, but existing methods solve the allocation for a single fixed memory budget.

By Ashitabh Misra, Madhav Agrawal, Arham Jain, Tarek Abdelzaher
llmsefficiency
More like this →
arXiv AI
Aug 6

Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs

arXiv:2608. 04048v1 Announce Type: cross Abstract: Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput.

By Yu Luo, Bo Dong, Wenhua Cheng, Haihao Shen
llmsefficiency
More like this →