arXiv Machine Learning By Eunju Shin, Jongbin Ryu

Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing

Read the original on arXiv Machine Learning →

The paper introduces Colla-Q, a Mixture-of-Experts (MoE) quantization technique that uses activation entropy to allocate bit-widths across experts. By balancing performance among experts, Colla-Q improves overall MoE accuracy and reduces reliance on calibration datasets. The method aims to maintain robustness and stability in quantized MoE models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 28

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets

GAMMA is a post‑training framework that learns module‑wise precision preferences for mixed‑precision quantization of large language models. It optimizes a teacher‑forced hidden‑state reconstruction objective under an augmented Lagrangian constraint and then projects the learned preferences into exact budget‑feasible discrete assignments via integer programming. Because the learned preferences encode a stable sensitivity ranking, a single training run can be reused for any deployment budget, reducing per‑budget adaptation from hours to minutes and outperforming fixed‑precision baselines and search‑based methods on Llama and Qwen models.

By Zhangyang Yao, Haiyan Zhao, Haoyu Wang, Xu Han
arXiv Machine Learning
Jun 2

WINDQuant: Weight-Informed Neural Decision-Making for Global Mixed-Precision LLM Quantization

arXiv:2605. 26660v2 Announce Type: replace Abstract: Quantization is an effective approach to reduce the memory footprint and inference cost of large language models (LLMs), yet maintaining performance in the ultra-low-bit regime remains challenging.

By Phong Nam Huu Nguyen, Khoi M. Le, Cong-Duy T Nguyen, Anh Tuan Luu, Thong Thanh Nguyen, Tho Quan
arXiv Machine Learning
Aug 27

FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference

The paper introduces FAMPWQ, a Fisher information-based Adaptive Mixed Precision Weight Quantization method designed to improve Large Language Model inference on commodity GPUs. It uses a Fisher information metric to assess layer-wise sensitivity and a reinforcement learning-based bit-width allocator to adaptively assign precision per layer. Experiments across seven models and five benchmarks show significant gains in perplexity, accuracy, and LLM-as-a-judge performance compared to seven baseline approaches.

By Gongwei Lee, Ji Liu, Juncheng Jia, Ji Wu