arXiv AI By Michael S. Zhang, Rishi A. Ruia, Arnav Kewalram, Saathvik Dharmapuram, Utkarsh Sharma, Kevin Zhu

When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models

Read the original on arXiv AI →

arXiv:2512. 18934v2 Announce Type: replace-cross Abstract: Catastrophic forgetting poses a fundamental challenge in continual learning, particularly when models are quantized for deployment efficiency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 31

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

The paper introduces DAMP, a decay‑aware mixed‑precision quantization scheme for recurrent‑state representations in GDN and KDA language models. By identifying high‑risk channels through quantization‑error energy and decay persistence, DAMP stores these channels at higher precision while compressing the rest to INT8, achieving a 9.9‑bit average precision. Experiments on Qwen3.6‑35B and Kimi‑Linear‑48B show a 69.1% reduction in recurrent‑state storage, up to 2.01× faster state‑update kernels, and up to 10.9% lower full‑model TPOT while preserving accuracy close to the FP32 baseline.

By Tao Zhang, Jianchao Tan, Pingwei Sun, Yanqi Yu, Zixu Jiang, Yuchen Xie, Xunliang Cai, Ziqian Zeng
arXiv Machine Learning
Aug 20

Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs

The study investigates how post‑training quantization (PTQ) affects proactive interference (PI) in large language models. Using bitsandbytes, the authors compare FP16, INT8, and INT4/NF4 precision across three instruction‑tuned models and find that INT4 quantization markedly degrades accuracy under high interference, with INT8 also incurring a smaller penalty in two of the three models. The degradation is linked to increased same‑key intrusion errors and originates in the quantized transformer backbone rather than the output layer.

By Shayan Shahrabi-Farahani (Shahid Beheshti University, Tehran, Iran), Dara Rahmati (Shahid Beheshti University, Tehran, Iran)
arXiv Machine Learning
Sep 11

Why Does Post-Training Quantization Work?

Post‑training quantization compresses large language models by storing weights at reduced precision, introducing errors into hidden states that could accumulate with depth. However, pretrained models accumulate far less hidden‑state error than randomly initialized ones, largely preserving downstream performance. The study identifies two key mechanisms: (1) each layer’s new error tends to oppose inherited error, partially canceling it, and (2) the LM‑head geometry preserves high‑rank token scores, mitigating output changes.

By Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen