Introducing AutoRound: Intel’s Advanced Quantization for LLMs and VLMs
Related stories
VQ-bench: A Composable Vector Quantization Framework
arXiv:2608. 11240v1 Announce Type: new Abstract: Vector quantization is an old problem but has recently become central to AI infrastructure.
A Target-Centric Survey of Quantization-Aware Training
The paper presents a target‑centric survey of Quantization‑Aware Training (QAT), a technique that simulates quantization during model training to produce low‑bit models with accuracy comparable to full‑precision ones. It systematically reviews existing QAT methods using a target‑centric taxonomy, highlighting differences in error characteristics, numerical formats, and strategy transferability across targets. The survey also summarizes QAT evaluation paradigms, discusses optimization and deployment challenges, and outlines potential future research directions.
ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization
arXiv:2608. 11045v1 Announce Type: new Abstract: ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals.
Making LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA
Rethinking Small VLM Quantization: From Component-Wise Analysis to Hardware-Aware Edge Deployment
arXiv:2607. 08029v1 Announce Type: new Abstract: The emergence of vision language models with fewer than 3 billion parameters has accelerated the implementation of on-device multimodal intelligence.
FPTQuant: Function-Preserving Transforms for LLM Quantization
arXiv:2506. 04985v2 Announce Type: replace Abstract: Large language models (LLMs) require substantial compute, and thus energy, at inference time.
MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration
MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration proposes a new quantization format that inverts the traditional microscaling approach by assigning private exponents to each element and a shared mantissa. The adaptive dual-format MiX-MX inference framework maps this format to a custom accelerator, replacing multipliers with shifters. Evaluations show that 4.5-bit MiX matches or surpasses NVFP4 accuracy on multimodal benchmarks while improving area efficiency by 25% and delivering 2.3–4.5× speedup with 1.4–2.9× energy reduction compared to the Focus accelerator.
Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
The paper introduces Llama-Mobile, a framework that quantizes vision‑language models for efficient mobile deployment. It uses a quantization pipeline that generates training data from the model itself, eliminating the need for the original training setup, and employs a novel 2.7‑bit‑per‑parameter format optimized for Arm CPUs. Applying this method, the authors compress the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8‑bit activations while maintaining strong performance on visual question answering tasks.
GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
arXiv:2607. 27694v1 Announce Type: cross Abstract: Low-bit quantization is essential for efficient LLM inference, and both rotation and fine-grained group quantization have shown individual promise.
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference
arXiv:2607. 17733v1 Announce Type: cross Abstract: 4-bit quantization enables efficient LLM inference, but suffers from significant accuracy degradation due to outliers.
HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference
HBQ: Hierarchical Scaling Block Quantization with Hardware‑Efficiency‑Aware Design for Accurate LLM Inference proposes a new block‑quantization scheme that uses large blocks and low‑overhead significand scaling to balance hardware efficiency and accuracy. The authors demonstrate that larger blocks improve efficiency by amortizing dequantization and accumulation costs, while their SIG scaling compensates for the resulting accuracy loss. Experiments on a 28 nm ASIC accelerator show that HBQ achieves up to 4.6× higher area/energy efficiency than state‑of‑the‑art weight‑only quantization, with 1.5–3.0× speedup and 1.6–3.3× system energy reduction over existing BQ methods.