arXiv Machine Learning

Leech Lattice Vector Quantization for Efficient LLM Compression

arXiv:2603. 11021v2 Announce Type: replace Abstract: Scalar quantization of large language models (LLMs) is fundamentally limited by information-theoretic bounds.