TORQUE: Optimizing What (not) to Quantize Before and After Rotation
Read the original on arXiv AI →The paper introduces TORQUE, a framework that enhances quantization by jointly optimizing which coordinates to keep at high precision before and after applying uniform random rotations, all within a fixed bit budget. By preserving large input coordinates before rotation and the largest-magnitude coordinates after rotation, TORQUE reduces quantization error and allows efficient use of offline-optimized codebooks. The authors provide an error upper bound, prove that top‑k pre‑rotation retention is optimal for each k, and demonstrate improved accuracy‑storage tradeoffs in Gaussian models and practical tasks such as nearest‑neighbor retrieval, KV‑cache compression, and activation compression.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.