arXiv AI

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels

arXiv:2606. 07713v1 Announce Type: cross Abstract: The attention mechanism is the dominant computational bottleneck in modern transformer-based AI.

arXiv Machine Learning
Sep 22

FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax Attention

FlashBoB introduces an I/O‑efficient algorithm for exact backward‑over‑backward (BoB) in softmax attention, enabling precise second‑order differentiation without large intermediate tensors. By exploiting a hierarchical affine structure, the method confines computation to on‑chip tiles and limits off‑chip memory traffic, achieving θ(N² d²/M) HBM usage. Experiments show FlashBoB scales to sequence lengths of 262K on a single A100 GPU, outperforming prior exact baselines and FlashBack by up to 6.3×.

By Anthony Givans, Michael Crawshaw, Mingrui Liu
arXiv Machine Learning
1d ago

Fast Polynomial Transcendentals for LLMs

The paper investigates using short polynomial approximations to accelerate special‑function operations in large language models on NVIDIA Blackwell GPUs. By replacing native sigmoid, tanh, and SiLU with degree‑3 or degree‑4 bfloat16 programs, the authors achieve up to 2.19× speed‑ups in isolated FP16 benchmarks and modest training‑step throughput gains (2.7–8.0%) across four integration tasks. The study also evaluates model behavior, finding negligible training‑loss differences within 100 billion tokens.

By Robert Hu
arXiv Machine Learning
Aug 4

Nova: An End-to-End MLIR Compiler for Deep Learning

arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.

By Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra
arXiv Machine Learning
Jun 9

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy

arXiv:2606. 09080v1 Announce Type: new Abstract: Pruning has emerged as a dominant paradigm for accelerating large language model (LLM) inference, spanning a broad spectrum of methods that remove computation across tokens, layers, heads, dimensions, and attention patterns.

By Haozhe Hu, Hao Wu, Anhao Zhao, Longwei Ding, Peiran Yin, Yunpu Ma, Xiaoyu Shen
arXiv Machine Learning
1d ago

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

The paper introduces TACO, a new optimizer for fine‑tuning large language models that drastically reduces optimizer state memory while preserving first‑order gradients. TACO selects the sign of the largest magnitude entry in each column of weight matrices, achieving a 174× reduction in persistent optimizer memory compared to AdamW8bit and a 2.9× decrease in peak training memory on OPT‑13B. This allows full‑parameter fine‑tuning of 30–32B‑parameter models on a single 80 GB GPU across multiple model families and tasks, with comparable accuracy and runtime to existing methods.

By Jichao Jiang (University of Central Florida), Cristian McGee (University of Central Florida), El Houcine Bergou (Mohammed VI Polytechnic University), Hanqin Cai (University of Central Florida), Aritra Dutta (University of Central Florida)
arXiv AI
Jun 19

Efficiently Representing Algorithms With Chain-of-Thought Transformers

arXiv:2606. 19697v1 Announce Type: cross Abstract: The increasing popularity of \emph{reasoning} models -- language models that output a series of reasoning or thought tokens before producing an answer -- is justified, in part, by theoretical results showing that chain-of-thought (CoT) transformers can simulate Turing machines, and thus perform arbitrary computation.

By Yanhong Li, Anej Svete, Ashish Sabharwal, William Merrill
arXiv Machine Learning
1d ago

TANGO: Treating Tokens as Operators

The paper introduces TANGO, a Token‑Aggregated Nonlinear Gating Operator that blends cross‑token mixing and token‑wise transformation in transformer architectures. By computing a nonlinear gate per source token and averaging these gates for each destination, TANGO forms a source‑conditioned linear operator that improves predictive performance. Experiments on web text, formal mathematics, and code show that full‑prefix TANGO achieves the lowest test negative log‑likelihood across 16 settings, while a narrower variant offers substantial throughput gains with only a modest increase in loss.

By Joshua Nunley
arXiv Machine Learning
5d ago

Benchmarking Attention for Tabular Foundation Models

The paper introduces a reproducible benchmark for evaluating attention mechanisms in tabular foundation models, focusing on the distinct row and column attention patterns that differ from language model attention. It compares several backends—Torch SDPA, FlashAttention variants, vLLM, and SageAttention—across realistic tabular shapes on A100, H100, and B200 GPUs, revealing that optimal backend choice varies by attention type, hardware, and model specifics. The study finds FlashAttention generally performs best, but CuDNN can outperform it for column attention on longer sequences, while SageAttention excels for large row sequences beyond 16k rows.

By Maximilian Schambach, Clemens Biehl, Sam Thelin