VQ-bench: A Composable Vector Quantization Framework
arXiv:2608. 11240v1 Announce Type: new Abstract: Vector quantization is an old problem but has recently become central to AI infrastructure.
Related stories
Inner Product Aware Quantization: Provably Fast, Accurate, and Adaptive Algorithms
arXiv:2606. 00289v1 Announce Type: new Abstract: Quantization is a fundamental tool used to compress datasets, neural network weights, and memory usage in a range of computational tasks.
A Target-Centric Survey of Quantization-Aware Training
The paper presents a target‑centric survey of Quantization‑Aware Training (QAT), a technique that simulates quantization during model training to produce low‑bit models with accuracy comparable to full‑precision ones. It systematically reviews existing QAT methods using a target‑centric taxonomy, highlighting differences in error characteristics, numerical formats, and strategy transferability across targets. The survey also summarizes QAT evaluation paradigms, discusses optimization and deployment challenges, and outlines potential future research directions.
LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization
arXiv:2606. 10531v1 Announce Type: cross Abstract: Quantization-aware training (QAT) is essential for extremely low-bit large language models (LLMs).
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
arXiv:2505. 18231v3 Announce Type: replace-cross Abstract: Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to the large size of key-value (KV) cache.
Introducing AutoRound: Intel’s Advanced Quantization for LLMs and VLMs
Efficient AI Model Deployment Using Quantization Analysis Tool
The paper introduces the Quantization Analysis Tool, a system built on the ONNX framework that streamlines quantization workflows for deep learning models. It offers layer‑wise sensitivity analysis, visualizations of weight and activation distributions, and guidance for selecting precision levels to balance model size, latency, and accuracy. Experiments on various neural network architectures show that the tool improves quantized accuracy and overall deployment efficiency.
FPTQuant: Function-Preserving Transforms for LLM Quantization
arXiv:2506. 04985v2 Announce Type: replace Abstract: Large language models (LLMs) require substantial compute, and thus energy, at inference time.
ADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language Models
arXiv:2605.11222v2 Announce Type: replace Abstract: Quantization is an effective strategy to reduce the storage and computation footprint of large language models (LLMs). Post-training quantization (...
WaterSIC: Information-Theoretically (Near) Optimal Linear Layer Quantization
arXiv:2603. 04956v2 Announce Type: replace Abstract: This paper considers the problem of converting a given dense linear layer to low precision.
ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization
arXiv:2608. 07019v1 Announce Type: new Abstract: Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models.
Joint Structural Pruning and Mixed-Precision Quantization for LLM Compression
arXiv:2606. 07819v1 Announce Type: new Abstract: Recently, the efficiency of Large Language Models (LLMs) deployment has become a critical concern in practical applications.