arXiv Machine Learning By Gunjun Lee, Sehwan Son, Younjoo Lee, Byungjun Kim, Jung Ho Ahn

SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization

Read the original on arXiv Machine Learning →

arXiv:2608. 15567v1 Announce Type: new Abstract: Weight-only post-training quantization (PTQ) enables the deployment of large language models under tight memory budgets, but accuracy often collapses at 2-3 bits.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.