Recent foundation models are moving toward native multimodal Vision-Language Models (VLMs), making VLMs a central form of next-generation foundation models. However, their large language backbones mak...
arXiv:2606. 15652v1 Announce Type: new Abstract: 4-bit quantization significantly reduces the memory footprint and accelerates the inference of large language models (LLMs).
By Yangjia Hu, Haodong Wang, Zicong Hong, Qianli Liu, Quanxin Shou, Jian Lin, Song Guo, Xiaowei Shen, Xiangjun Huang, Dian Wang, Jian Yang
arXiv:2507. 23035v4 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated impressive capabilities across a wide range of applications, but demand substantial memory and compute resources during inference.
By Xueying Wu, Baijun Zhou, Zhihui Gao, Yuzhe Fu, Qilin Zheng, Yintao He, Hai Li
arXiv:2502. 00527v2 Announce Type: replace Abstract: The KV cache in large language models is a dominant factor in memory usage, limiting their broader applicability.
By Songhao Wu, Ang Lv, Xiao Feng, Yufei Zhang, Xun Zhang, Guojun Yin, Wei Lin, Rui Yan
arXiv:2608.30384v1 Announce Type: new
Abstract: By introducing RSLM (Rotated Scaled Lloyd-Max), a family of training-free vector quantization codecs compressing embeddings to 1--4 bits per dimension,...
By Rastislav Lenhardt, Teodora Dobos, Thomas Vecchiato, Jiri Isa, Igor Ginzburg
The paper introduces TORQUE, a framework that enhances quantization by jointly optimizing which coordinates to keep at high precision before and after applying uniform random rotations, all within a fixed bit budget. By preserving large input coordinates before rotation and the largest-magnitude coordinates after rotation, TORQUE reduces quantization error and allows efficient use of offline-optimized codebooks. The authors provide an error upper bound, prove that top‑k pre‑rotation retention is optimal for each k, and demonstrate improved accuracy‑storage tradeoffs in Gaussian models and practical tasks such as nearest‑neighbor retrieval, KV‑cache compression, and activation compression.
By Ran Ben Basat, Michael Mitzenmacher, Shay Vargaftik