arXiv Machine Learning By Joshua Hill

Saturation Makes Quantization Error Additive: A Coverage Model with a Certificate

Read the original on arXiv Machine Learning →

arXiv:2607. 12266v1 Announce Type: new Abstract: Mixed-precision quantization must decide which parts of a model to keep at higher precision.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 2

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

The paper investigates where post‑training quantization (PTQ) harms large language models (LLMs) and how to best allocate a limited precision budget. By causally raising each layer to 8‑bit precision across nine open‑weight models, the authors find that quantization damage is diffuse rather than concentrated in specific task circuits or weight statistics, and that globally refining quantization granularity outperforms selectively protecting the most recoverable layers. They also observe that the residual accuracy loss is budget‑limited and that peak recovery locations correlate with architecture within families but not across families.

By Jundong Hu, Shekar Ramachandran
arXiv Machine Learning
Sep 24

Predicting Quantization Price for Selecting PTQ Configurations Before Deployment

The paper proposes a method for selecting post‑training quantization (PTQ) configurations before deployment by treating each admissible layer configuration as an error generator with an associated deployment cost. It introduces a priced layer‑output error framework that uses the covariance of layer outputs and the full‑precision model’s curvature to compute a price for each configuration. This approach replaces traditional reconstruction or Hessian‑based scores with a unified, cost‑aware selector that can calibrate and budget PTQ settings efficiently.

By Junbin Qiu, Jian Mu, Weitong Zhang, Yao Shu