arXiv Machine Learning

A Quantum Roadmap for Softmax Attention: Exact Born-Rule Analogs for Softmax Attention on the Probability Simplex

arXiv:2608. 11173v1 Announce Type: cross Abstract: The attention mechanism forms the foundation of many modern AI models such as the Transformer.

arXiv Machine Learning
Sep 21

EFQ-Softmax: Exp-Free Quantization for Softmax

EFQ-Softmax is a low‑bit probability‑generation technique that replaces the traditional exp‑then‑quantize path in Transformer attention. It maps shifted attention scores directly to block‑scaled E2M1 operands using an exponent‑only scale and a single affine rule, allowing the same low‑bit representation to be used for both numerator and denominator updates. Experiments on Qwen3‑8B, Qwen3‑VL‑8B‑Instruct, and WAN2.2‑TI2V‑5B show that EFQ‑Softmax maintains or improves model quality while reducing vector‑stage latency by about 40% on the A5 vector unit.

By Haohui Han (Xi'an Jiaotong University), Yuming Wan (Huawei Technologies Co., Ltd), Hongni Wang (Shandong University of Finance and Economics), Pengcheng Xie (Huawei Technologies Co., Ltd), Xiaodong Yan (Xi'an Jiaotong University), Runqi You (Xi'an Jiaotong University), Wencong Zhang (Xi'an Jiaotong University)
arXiv Machine Learning
Sep 4

Entropy-Generated Attention Beyond Softmax and Entmax: Kaniadakis and Reciprocal-Symmetric Abe Operators

The paper introduces two novel attention operators derived from generalized statistical entropies. The Kaniadakis entropy yields a full-support normalization with algebraically decaying weights, while the Abe entropy produces an implicit reciprocal-symmetric operator. The authors analyze these operators through a Fisher-metric Lagrangian framework, compare them to Softmax and entmax, and provide a tangent-gradient test to distinguish changes in attention profiles from mere scaling effects.

By Gunn Kim
arXiv Machine Learning
Sep 21

Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

The paper investigates how gating the value pathway in attention mechanisms provides two missing capabilities of softmax attention: abstention and noise filtering. Experiments on models ranging from 10M to 350M parameters show that abstention benefits smaller models while noise filtering becomes more advantageous as models scale, and that combining both primitives yields the best performance across all sizes. The authors also demonstrate that the gates effectively suppress interference and that each gate type has a distinct blind spot, all while adding negligible parameters and preserving compatibility with key‑value caching.

By Richard Zhe Wang