arXiv Machine Learning

SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization

arXiv:2608. 15567v1 Announce Type: new Abstract: Weight-only post-training quantization (PTQ) enables the deployment of large language models under tight memory budgets, but accuracy often collapses at 2-3 bits.