arXiv AI By Jianlin Yu, Jing Lin, Linghui Kong, Aiyue Chen, Weiyi Sun, Chenyu Zeng, Wangli Lan, Jinxi Li, Zhuo Zheng, Ziyang Yue, Danning Ke, Fei Yi, Tianchi Hu, Yuan Ding, Yiwu Yao, Junsong Wang

MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention

Read the original on arXiv AI →

arXiv:2607. 24377v1 Announce Type: cross Abstract: The quadratic cost of attention is a major bottleneck in diffusion-based video generation models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 21

EFQ-Softmax: Exp-Free Quantization for Softmax

EFQ-Softmax is a low‑bit probability‑generation technique that replaces the traditional exp‑then‑quantize path in Transformer attention. It maps shifted attention scores directly to block‑scaled E2M1 operands using an exponent‑only scale and a single affine rule, allowing the same low‑bit representation to be used for both numerator and denominator updates. Experiments on Qwen3‑8B, Qwen3‑VL‑8B‑Instruct, and WAN2.2‑TI2V‑5B show that EFQ‑Softmax maintains or improves model quality while reducing vector‑stage latency by about 40% on the A5 vector unit.

By Haohui Han (Xi'an Jiaotong University), Yuming Wan (Huawei Technologies Co., Ltd), Hongni Wang (Shandong University of Finance and Economics), Pengcheng Xie (Huawei Technologies Co., Ltd), Xiaodong Yan (Xi'an Jiaotong University), Runqi You (Xi'an Jiaotong University), Wencong Zhang (Xi'an Jiaotong University)
arXiv Machine Learning
Aug 28

Activation Outliers Matter: Robust Recovery for Quantized Multimodal LLMs

The paper investigates low‑bit quantization for Multimodal Large Language Models (MLLMs), showing that MXFP8 retains near‑lossless performance while 4‑bit formats like MXFP4 and HiF4 cause significant degradation. It identifies activation quantization as the main source of this loss and introduces Residual Fallback Quantization (RFQ), a lightweight framework that adds a quantized residual pathway to improve activation fidelity without architectural changes. Experiments on Wan2.2 and Qwen3‑VL demonstrate that RFQ recovers much of the performance gap to BF16 baselines across generation and reasoning tasks.

By Tanzila Rahman, Mehran Taghian Jazi, Yunke Peng, Zhuang Ma, Anandharaju Durai Raju, Yao Wang, Xing Huang, Hei Yi Mak, Shadan Golestan, Hoang Le, Yonghan Dong, Wei Guo, Yaoyuan Wang