arXiv Machine Learning By Miguel P. Bento, Jo\~ao Seabra

GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries

Read the original on arXiv Machine Learning →

arXiv:2607. 20757v1 Announce Type: new Abstract: Transformers are known to have internal continuous symmetries that leave outputs invariant, while modifying quantization.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
4d ago

ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs

ThinQuant introduces efficient rotation learning for low‑bit weight and activation quantization of large language models by reducing calibration data through a geometric selection of activations and solving a lower‑dimensional optimization problem via an ADMM algorithm. The method achieves comparable or better quantization performance with dramatically fewer calibration points, completing rotation calibration for Llama‑3‑70B in under 12 minutes and for Llama‑3.1‑405B in just over 2 hours on a single GPU. ThinQuant outperforms existing gradient‑free approaches such as DartQuant and gradient‑based SpinQuant in both speed and perplexity metrics on WikiText‑2.

By Mehdi Makni, Ryan Lucas, Rahul Mazumder
arXiv Machine Learning
1d ago

ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs

ConQuR introduces a lightweight post‑training rotation calibration for large language model activation quantization. By learning orthogonal rotations that align normalized activations with the corners of an inscribed hypercube, the method distributes activation energy evenly and can be updated online without storing activations. Experiments on Llama‑2 and Llama‑3 models (3B–70B) show competitive or improved perplexity and reasoning performance while avoiding costly training or large offline storage.

By Chayne Thrash, Ali Abbasi, Soheil Kolouri