FlashBoB introduces an I/O‑efficient algorithm for exact backward‑over‑backward (BoB) in softmax attention, enabling precise second‑order differentiation without large intermediate tensors. By exploiting a hierarchical affine structure, the method confines computation to on‑chip tiles and limits off‑chip memory traffic, achieving θ(N² d²/M) HBM usage. Experiments show FlashBoB scales to sequence lengths of 262K on a single A100 GPU, outperforming prior exact baselines and FlashBack by up to 6.3×.
By Anthony Givans, Michael Crawshaw, Mingrui Liu
arXiv:2607. 19456v1 Announce Type: cross Abstract: We derive four memory-optimal inference artifacts for transformer attention using the Mathematics of Arrays (MoA), each following directly from the forward-pass Denotational Normal Form (DNF) of with the query-row index fixed to the current decode step.
By Lenore Mulin, Gaetan Hains
arXiv:2609.13612v1 Announce Type: new
Abstract: Modern AI systems are built on the Transformer architecture, whose core operation, attention, accounts for the majority of computation and memory cost....
By Varun Kumar Dasoju, Tian Zhao
The paper investigates using short polynomial approximations to accelerate special‑function operations in large language models on NVIDIA Blackwell GPUs. By replacing native sigmoid, tanh, and SiLU with degree‑3 or degree‑4 bfloat16 programs, the authors achieve up to 2.19× speed‑ups in isolated FP16 benchmarks and modest training‑step throughput gains (2.7–8.0%) across four integration tasks. The study also evaluates model behavior, finding negligible training‑loss differences within 100 billion tokens.
By Robert Hu
arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.
By Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra
arXiv:2606. 09080v1 Announce Type: new Abstract: Pruning has emerged as a dominant paradigm for accelerating large language model (LLM) inference, spanning a broad spectrum of methods that remove computation across tokens, layers, heads, dimensions, and attention patterns.
By Haozhe Hu, Hao Wu, Anhao Zhao, Longwei Ding, Peiran Yin, Yunpu Ma, Xiaoyu Shen