arXiv Machine Learning

Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization

arXiv:2606. 02288v1 Announce Type: new Abstract: Massive activation spikes in Large Language Models (LLMs) severely degrade quantization by stretching dynamic ranges.

Hugging Face Trending Papers
Jun 3

STaR-Quant: State-Time Consistent Post-Training Quantization for Diffusion Large Language Models

Diffusion large language models (DLLMs) have recently emerged as a promising alternative to autoregressive LLMs by generating text through iterative masked denoising with bidirectional context. However, their large model sizes and iterative denoising process introduce substantial memory and computational overhead, motivating post-training quantization for efficient deployment.

arXiv Machine Learning
1d ago

ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs

ConQuR introduces a lightweight post‑training rotation calibration for large language model activation quantization. By learning orthogonal rotations that align normalized activations with the corners of an inscribed hypercube, the method distributes activation energy evenly and can be updated online without storing activations. Experiments on Llama‑2 and Llama‑3 models (3B–70B) show competitive or improved perplexity and reasoning performance while avoiding costly training or large offline storage.

By Chayne Thrash, Ali Abbasi, Soheil Kolouri
arXiv Computation and Language
Aug 25

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

arXiv:2608.12149v2 Announce Type: replace Abstract: We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies:...

By Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo