arXiv Machine Learning

Predicting Quantization Price for Selecting PTQ Configurations Before Deployment

The paper proposes a method for selecting post‑training quantization (PTQ) configurations before deployment by treating each admissible layer configuration as an error generator with an associated deployment cost. It introduces a priced layer‑output error framework that uses the covariance of layer outputs and the full‑precision model’s curvature to compute a price for each configuration. This approach replaces traditional reconstruction or Hessian‑based scores with a unified, cost‑aware selector that can calibrate and budget PTQ settings efficiently.

arXiv AI
Sep 2

REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

REAL-Q introduces a new post‑training quantization approach for large language models that replaces the traditional single closed‑form second‑order solver with a fine‑grained, dynamic block‑wise gradient descent applied after every 128‑column block. By aligning the surrogate loss with the end‑to‑end objective and using a sliding window for smooth cross‑layer transitions, REAL‑Q mitigates error propagation and information misalignment. Experiments on LLaMA‑3.1 and Qwen3 show up to ~49% reduction in end‑to‑end KL divergence compared to state‑of‑the‑art methods.

By Qian Zhang, Yaoming Li, Zhewen Tan, Yanshu Wang, Heng Lu, Kun Su, Zongwei Lv, Wenhan Yu, Yongge Ma, Yinjun Han, Ruikuang Liu, Tong Yang