arXiv Machine Learning

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs

arXiv:2602. 10431v4 Announce Type: replace Abstract: Large language models (LLMs) demand substantial computational and memory resources, posing challenges for efficient deployment.

arXiv Machine Learning
Sep 1

A Target-Centric Survey of Quantization-Aware Training

The paper presents a target‑centric survey of Quantization‑Aware Training (QAT), a technique that simulates quantization during model training to produce low‑bit models with accuracy comparable to full‑precision ones. It systematically reviews existing QAT methods using a target‑centric taxonomy, highlighting differences in error characteristics, numerical formats, and strategy transferability across targets. The survey also summarizes QAT evaluation paradigms, discusses optimization and deployment challenges, and outlines potential future research directions.

By Jiamin Song, Mengjie Zhao, Zijing Wang, Yongkang Liu, Qian Li, Shi Feng, Feiliang Ren, Daling Wang, Hinrich Sch\"utze
arXiv Machine Learning
Jun 2

WINDQuant: Weight-Informed Neural Decision-Making for Global Mixed-Precision LLM Quantization

arXiv:2605. 26660v2 Announce Type: replace Abstract: Quantization is an effective approach to reduce the memory footprint and inference cost of large language models (LLMs), yet maintaining performance in the ultra-low-bit regime remains challenging.

By Phong Nam Huu Nguyen, Khoi M. Le, Cong-Duy T Nguyen, Anh Tuan Luu, Thong Thanh Nguyen, Tho Quan
arXiv Machine Learning
Sep 21

SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference

SpecQuant is a training‑free framework that merges speculative decoding with multi‑parent quantization to enable adaptive, efficient inference of large language models. It generates several quantized variants (INT4, FP8, FP16) from a single base model and routes queries to the appropriate variant based on predicted complexity, using lightweight models for simple tasks and full‑precision models for complex reasoning. Evaluations on Qwen2.5 models across MMLU, AlpacaEval, and GSM8K show 35‑43% speedups with less than 2% accuracy loss, facilitating practical on‑device LLM deployment without specialized infrastructure.

By Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T
arXiv Machine Learning
Sep 10

Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method

Squeeze10-LLM is a staged mixed‑precision post‑training quantization framework that reduces 16‑bit LLM weights to an average of 1.6 bits per weight by assigning 80% of weights to 1 bit and 20% to 4 bits. It introduces Post‑Binarization Activation Robustness (PBAR), a weight significance metric that considers activation impact, and Full Information Activation Supervision (FIAS), a strategy that preserves activation information to limit error propagation. Experiments on LLaMA and LLaMA2 demonstrate that Squeeze10‑LLM achieves state‑of‑the‑art performance for sub‑2‑bit weight‑only quantization, raising average accuracy from 43% to 56% on six zero‑shot classification tasks.

By Qingcheng Zhu, Yangyang Ren, Linlin Yang, Yanjing Li, Sheng Xu, Haodong Zhu, Juan Zhang, Runqi Wang, Baochang Zhang