arXiv AI By Ran Ben Basat, Michael Mitzenmacher, Shay Vargaftik

TORQUE: Optimizing What (not) to Quantize Before and After Rotation

Read the original on arXiv AI →

The paper introduces TORQUE, a framework that enhances quantization by jointly optimizing which coordinates to keep at high precision before and after applying uniform random rotations, all within a fixed bit budget. By preserving large input coordinates before rotation and the largest-magnitude coordinates after rotation, TORQUE reduces quantization error and allows efficient use of offline-optimized codebooks. The authors provide an error upper bound, prove that top‑k pre‑rotation retention is optimal for each k, and demonstrate improved accuracy‑storage tradeoffs in Gaussian models and practical tasks such as nearest‑neighbor retrieval, KV‑cache compression, and activation compression.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 22

SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models

SPHQuant introduces a rotation‑free spherical weight‑only quantization framework for Vision‑Language Models, decomposing 8‑dimensional weight vectors into sign, radius, and a positive unit direction. By isolating outlier magnitudes in the radius and allocating extra precision there, it mitigates accuracy loss at extreme low bit‑widths. The method also employs a compact positive‑direction codebook with angular fine‑tuning and a hardware‑friendly GEMV kernel, achieving state‑of‑the‑art performance while boosting decode throughput by 30.3% on RTX A6000 compared to QTIP.

By Kewei Zhang, Zheng Chen, Haotong Qin, Yulun Zhang