arXiv AI By Hanzhi Zhang, Qiao Zhang, Qinglei Cao, Heng Fan, Yan Huang, Kewei Sha, Yunhe Feng

AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Aug 19

TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration

TileMix is a tile‑centric mixed‑precision attention kernel that routes score‑tile groups within fused dense attention to either FP16 or INT8 computation, using compact bitmasks to decide precision per tile. By partitioning the attention matrix into hardware‑aligned tiles and updating a shared online‑softmax state, TileMix preserves dense token connectivity without requiring training and supports grouped‑query attention, variable‑length batches, and INT8 key/value caches. Benchmarks on LLaMA, Qwen, and Vicuna show that TileMix restores long‑context quality lost with uniform INT8 and improves prefill throughput over FP16, offering a controllable accuracy‑efficiency trade‑off across model families.

By Hanzhi Zhang, Qiao Zhang, Qinglei Cao, Heng Fan, Yan Huang, Kewei Sha, Yunhe Feng
arXiv Machine Learning
Aug 11

RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention

arXiv:2608. 08081v1 Announce Type: cross Abstract: Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows linearly with context, and dozens of expert sublayers that must be paged on demand.

By Anthony. Lui, Mohamed. Elsaied, N. P. Savani