arXiv:2608. 14691v1 Announce Type: new Abstract: Sequence models are conventionally distinguished by their backbone, the mechanism that routes information across positions, such as attention or recurrence.
By Ahmed Nebli, Hadi Saadatdoorabi, Christopher Keibel, Kevin Yam
Low-bit attention accelerates Transformer inference by moving the $QK^\top$ and $PV$ matrix multiplications to FP8 or FP4 matrix engines. However, the softmax path often evaluates shifted-score expone...
arXiv:2606. 12059v1 Announce Type: new Abstract: We address transformer attention on energy-constrained physical substrates.
By Fabio Pasqualetti, Taosha Guo
arXiv:2512. 11784v2 Announce Type: replace Abstract: Softmax attention is a central component of transformer architectures, yet its nonlinear structure poses significant challenges for theoretical analysis.
By Etienne Boursier, Claire Boyer
EFQ-Softmax is a low‑bit probability‑generation technique that replaces the traditional exp‑then‑quantize path in Transformer attention. It maps shifted attention scores directly to block‑scaled E2M1 operands using an exponent‑only scale and a single affine rule, allowing the same low‑bit representation to be used for both numerator and denominator updates. Experiments on Qwen3‑8B, Qwen3‑VL‑8B‑Instruct, and WAN2.2‑TI2V‑5B show that EFQ‑Softmax maintains or improves model quality while reducing vector‑stage latency by about 40% on the A5 vector unit.
By Haohui Han (Xi'an Jiaotong University), Yuming Wan (Huawei Technologies Co., Ltd), Hongni Wang (Shandong University of Finance and Economics), Pengcheng Xie (Huawei Technologies Co., Ltd), Xiaodong Yan (Xi'an Jiaotong University), Runqi You (Xi'an Jiaotong University), Wencong Zhang (Xi'an Jiaotong University)
arXiv:2608. 09558v1 Announce Type: new Abstract: How expressive is prompting a transformer?
By Alexander Hsu, Rongjie Lai