arXiv AI By Mehdi Makni, Ryan Lucas, Rahul Mazumder

ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs

Read the original on arXiv AI →

ThinQuant introduces efficient rotation learning for low‑bit weight and activation quantization of large language models by reducing calibration data through a geometric selection of activations and solving a lower‑dimensional optimization problem via an ADMM algorithm. The method achieves comparable or better quantization performance with dramatically fewer calibration points, completing rotation calibration for Llama‑3‑70B in under 12 minutes and for Llama‑3.1‑405B in just over 2 hours on a single GPU. ThinQuant outperforms existing gradient‑free approaches such as DartQuant and gradient‑based SpinQuant in both speed and perplexity metrics on WikiText‑2.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization

HARP (Hadamard‑Preconditioned Adaptive Rotation Processor) is a learnable, structured two‑sided orthogonal processor that replaces fixed randomized Hadamard transforms in post‑training quantization of large language models. By representing rotations as sparse butterfly‑like block‑orthogonal stages and supporting mixed‑radix schedules, HARP adapts the quantization basis to each layer and calibration distribution while maintaining full‑precision equivalence. Across 2–4‑bit settings on Llama models from 1B to 70B, HARP consistently improves perplexity, delivers the strongest zero‑shot gains at 2 bits, and preserves deployment efficiency—achieving 128 tokens per second on Llama 2 7B at 2 bits, roughly 90% of RHT throughput and over twice the speed of FP16.

By Artur Zagitov, Gleb Molodtsov, Aleksandr Beznosikov
arXiv Machine Learning
Jul 24

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

arXiv:2607. 21446v1 Announce Type: new Abstract: Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent.

By Yann Bouquet, Alireza Khodamoradi, Kristof Denolf, Mathieu Salzmann