arXiv AI
Sep 17

HyQuant: Hybrid-Precision Quantization for LLM Attention

HyQuant introduces a hybrid-precision quantization framework for large language model (LLM) attention modules. It quantizes most attention states to low-bit formats while preserving a small set of vertical‑line tokens and local‑window states in full precision, guided by lightweight attention‑pattern signals. This design achieves near‑lossless accuracy across tasks while improving memory and hardware efficiency.

By Jiatong Ding, Bingxin Xing, Yu Zhang, Dian Ding, Xiaodong Yi, Xianbin Ouyang, Feihu Zhou, Kun Zhang, Zhenyu Guo, Hao Pan, Guangtao Xue, Yiming Zhang
arXiv AI
4d ago

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

arXiv:2609.38166v1 Announce Type: cross Abstract: Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Atte...

By Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica