Making LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA
Related stories
Introducing AutoRound: Intel’s Advanced Quantization for LLMs and VLMs
Understanding and Implementing Qwen3 From Scratch
A Detailed Look at One of the Leading Open-Source LLMs
LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection
arXiv:2606. 04050v1 Announce Type: cross Abstract: Existing quantization methods are fundamentally limited by rigid, integer-based bit-widths (e.
CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs
arXiv:2606. 26650v1 Announce Type: cross Abstract: In this paper, we present CAT-Q, Cost-efficient and Accurate Ternary Quantization, for compressing and accelerating LLMs.
CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs
In this paper, we present CAT-Q, Cost-efficient and Accurate Ternary Quantization, for compressing and accelerating LLMs. Unlike existing state-of-the-art ternary quantization methods that rely on data-intensive and costly quantization-aware training to mitigate severe performance degradation, CAT-Q is a simple yet effective post-training quantization scheme that is readily applicable to LLMs with diverse architectures and model sizes.
A Target-Centric Survey of Quantization-Aware Training
The paper presents a target‑centric survey of Quantization‑Aware Training (QAT), a technique that simulates quantization during model training to produce low‑bit models with accuracy comparable to full‑precision ones. It systematically reviews existing QAT methods using a target‑centric taxonomy, highlighting differences in error characteristics, numerical formats, and strategy transferability across targets. The survey also summarizes QAT evaluation paradigms, discusses optimization and deployment challenges, and outlines potential future research directions.
Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs
arXiv:2608. 04048v1 Announce Type: cross Abstract: Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput.
Channel-Wise Mixed-Precision Quantization for Large Language Models
arXiv:2410. 13056v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable success across a wide range of language tasks, but their deployment on edge devices remains challenging due to the substantial memory requirements imposed by their large parameter sizes.
Q-Strata: Hierarchical Bit Allocation for Mixed-Precision Quantization of Mixture-of-Experts LLMs
arXiv:2608.30564v1 Announce Type: cross Abstract: Mixed-precision quantization (MPQ) assigns a different bitwidth to each linear layer of a large language model (LLM) to minimize the quantization-ind...
Breaking the 1.58-bit Barrier for Ternary LLMs
arXiv:2609.16338v1 Announce Type: new Abstract: Ternary Large Language Models (LLM) store every weight as one of three symbols $\{-1,0,+1\}$, so the cost of a ternary model is conventionally referenc...
