arXiv AI By Yongge Ma, Guoan Wang, Feiyu Wang, Yaoming Li, Qian Zhang, Zihan Yan, Yinjun Han, Tong Yang

ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization

Read the original on arXiv AI →

arXiv:2608. 07019v1 Announce Type: new Abstract: Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.