arXiv Machine Learning
4d ago

Low-Discrepancy Dither for Quantized Recurrent State Caches

The paper investigates rounding strategies for low‑precision recurrent state caches in Mamba‑style and hybrid language models. It shows that a deterministic golden‑ratio Weyl dither consistently yields quantized models closer to full precision than stochastic rounding, across various models, storage formats, and long decoding horizons, without extra cost. In contrast, round‑to‑nearest can appear effective in short tests but degrades over long generations due to error accumulation.

By Snigdha Chandan Khilar
arXiv Machine Learning
Aug 31

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

The paper introduces DAMP, a decay‑aware mixed‑precision quantization scheme for recurrent‑state representations in GDN and KDA language models. By identifying high‑risk channels through quantization‑error energy and decay persistence, DAMP stores these channels at higher precision while compressing the rest to INT8, achieving a 9.9‑bit average precision. Experiments on Qwen3.6‑35B and Kimi‑Linear‑48B show a 69.1% reduction in recurrent‑state storage, up to 2.01× faster state‑update kernels, and up to 10.9% lower full‑model TPOT while preserving accuracy close to the FP32 baseline.

By Tao Zhang, Jianchao Tan, Pingwei Sun, Yanqi Yu, Zixu Jiang, Yuchen Xie, Xunliang Cai, Ziqian Zeng