ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM Weights
Read the original on arXiv AI →ShamAN-Q is a sub‑1‑bit post‑training quantization technique that builds on NanoQuant by replacing its diagonal reconstruction geometry with a dense curvature metric inspired by the Shampoo optimizer. For each linear weight, it fits a Kronecker product to the empirical Fisher information matrix of a small calibration set via Kullback–Leibler minimization, yielding a Mahalanobis reconstruction loss. The method updates continuous ADMM steps to Sylvester equations while keeping the discrete projection and deployment format unchanged, and it redistributes uniform rank across layers, achieving lower perplexity on Qwen3‑Base at roughly 1 bpw and matching or improving zero‑shot accuracy on the Eleuther LM Evaluation Harness.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.