arXiv Machine Learning

Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing

The paper introduces Colla-Q, a Mixture-of-Experts (MoE) quantization technique that uses activation entropy to allocate bit-widths across experts. By balancing performance among experts, Colla-Q improves overall MoE accuracy and reduces reliance on calibration datasets. The method aims to maintain robustness and stability in quantized MoE models.

arXiv AI
Aug 28

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets

GAMMA is a post‑training framework that learns module‑wise precision preferences for mixed‑precision quantization of large language models. It optimizes a teacher‑forced hidden‑state reconstruction objective under an augmented Lagrangian constraint and then projects the learned preferences into exact budget‑feasible discrete assignments via integer programming. Because the learned preferences encode a stable sensitivity ranking, a single training run can be reused for any deployment budget, reducing per‑budget adaptation from hours to minutes and outperforming fixed‑precision baselines and search‑based methods on Llama and Qwen models.

By Zhangyang Yao, Haiyan Zhao, Haoyu Wang, Xu Han
arXiv Machine Learning
Jun 2

WINDQuant: Weight-Informed Neural Decision-Making for Global Mixed-Precision LLM Quantization

arXiv:2605. 26660v2 Announce Type: replace Abstract: Quantization is an effective approach to reduce the memory footprint and inference cost of large language models (LLMs), yet maintaining performance in the ultra-low-bit regime remains challenging.

By Phong Nam Huu Nguyen, Khoi M. Le, Cong-Duy T Nguyen, Anh Tuan Luu, Thong Thanh Nguyen, Tho Quan
arXiv Machine Learning
Aug 27

FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference

The paper introduces FAMPWQ, a Fisher information-based Adaptive Mixed Precision Weight Quantization method designed to improve Large Language Model inference on commodity GPUs. It uses a Fisher information metric to assess layer-wise sensitivity and a reinforcement learning-based bit-width allocator to adaptively assign precision per layer. Experiments across seven models and five benchmarks show significant gains in perplexity, accuracy, and LLM-as-a-judge performance compared to seven baseline approaches.

By Gongwei Lee, Ji Liu, Juncheng Jia, Ji Wu
arXiv AI
Sep 1

Uncertainty Makes It Stable: Curiosity-Driven Quantized Mixture-of-Experts

The paper introduces a curiosity‑driven quantized Mixture‑of‑Experts framework that routes inputs based on Bayesian epistemic uncertainty across heterogeneous experts (BitNet ternary, 1‑16 bit BitLinear, post‑training quantization). On audio classification benchmarks, 4‑bit quantization preserves 99.9 % of full‑precision F1 while achieving 4× compression and 31 % energy savings, and curiosity‑driven routing further improves accuracy and reduces cross‑fold variance by up to 85 %. The routing is self‑organizing, allocating the most uncertain samples to the high‑precision expert, and the method demonstrates statistical parity with full precision across datasets.

By Sebasti\'an Andr\'es Cajas Ord\'o\~nez, Luis Fernando Torres Torres, Mackenzie J. Meni, Carlos Andr\'es Duran Paredes, Eric Arazo, Cristian Bosch, Ricardo Simon Carbajo, Yuan Lai, Leo Anthony Celi
arXiv AI
Jun 4

Model-Preserving Adaptive Rounding

arXiv:2505. 22988v3 Announce Type: replace-cross Abstract: The goal of quantization is to produce a compressed model whose output distribution is as close to the original model's as possible.

By Albert Tseng, Zhaofeng Sun, Christopher De Sa
arXiv AI
Jun 4

dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats

arXiv:2606. 04115v1 Announce Type: cross Abstract: Quantizing large language models (LLMs) to low-precision floating-point representations is central to efficient deployment, yet applying a single bit-width uniformly across all layers is sub-optimal in terms of both performance and accuracy.

By Giuseppe Franco, Ian Colbert, Pablo Monteagudo-Lago, Felix Marty, Nicholas Fraser
arXiv Machine Learning
Aug 11

RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention

arXiv:2608. 08081v1 Announce Type: cross Abstract: Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows linearly with context, and dozens of expert sublayers that must be paged on demand.

By Anthony. Lui, Mohamed. Elsaied, N. P. Savani