arXiv Machine Learning By Eric A. F. Reinhardt, Adam J. Hauser

A Quantum Roadmap for Softmax Attention: Exact Born-Rule Analogs for Softmax Attention on the Probability Simplex

Read the original on arXiv Machine Learning →

arXiv:2608. 11173v1 Announce Type: cross Abstract: The attention mechanism forms the foundation of many modern AI models such as the Transformer.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 21

EFQ-Softmax: Exp-Free Quantization for Softmax

EFQ-Softmax is a low‑bit probability‑generation technique that replaces the traditional exp‑then‑quantize path in Transformer attention. It maps shifted attention scores directly to block‑scaled E2M1 operands using an exponent‑only scale and a single affine rule, allowing the same low‑bit representation to be used for both numerator and denominator updates. Experiments on Qwen3‑8B, Qwen3‑VL‑8B‑Instruct, and WAN2.2‑TI2V‑5B show that EFQ‑Softmax maintains or improves model quality while reducing vector‑stage latency by about 40% on the A5 vector unit.

By Haohui Han (Xi'an Jiaotong University), Yuming Wan (Huawei Technologies Co., Ltd), Hongni Wang (Shandong University of Finance and Economics), Pengcheng Xie (Huawei Technologies Co., Ltd), Xiaodong Yan (Xi'an Jiaotong University), Runqi You (Xi'an Jiaotong University), Wencong Zhang (Xi'an Jiaotong University)