arXiv Machine Learning

Fast Gauss Sums via Flash Attention

The paper introduces a method to compute Gaussian kernel sums—central to many kernel-based machine learning techniques—using flash attention, a highly optimized softmax attention implementation. By applying two small input augmentations, the authors transform the normalized softmax reduction into an unnormalized Gauss sum, eliminating the need for custom GPU code. For feature dimensions greater than eight in fp16, this approach outperforms both compiled PyTorch code and PyKeOps kernels in speed, memory overhead, and accuracy, while maintaining linear memory scaling.

arXiv Machine Learning
Sep 22

FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax Attention

FlashBoB introduces an I/O‑efficient algorithm for exact backward‑over‑backward (BoB) in softmax attention, enabling precise second‑order differentiation without large intermediate tensors. By exploiting a hierarchical affine structure, the method confines computation to on‑chip tiles and limits off‑chip memory traffic, achieving θ(N² d²/M) HBM usage. Experiments show FlashBoB scales to sequence lengths of 262K on a single A100 GPU, outperforming prior exact baselines and FlashBack by up to 6.3×.

By Anthony Givans, Michael Crawshaw, Mingrui Liu
arXiv AI
Jun 19

StreamKL: Fast and Memory-Efficient KL Divergence for Boosting Attention Distillation

arXiv:2606. 20005v1 Announce Type: cross Abstract: Attention distillation, which trains one attention distribution to match another by minimizing their Kullback-Leibler (KL) divergence, is widely used in knowledge distillation, model compression, continual learning, and sparse-attention LLM training.

By Guangda Liu, Yiquan Wang, Chengwei Li, Wenhao Chen, Jing Lin, Yiwu Yao, Danning Ke, Wenchao Ding, Jieru Zhao
arXiv Machine Learning
1d ago

Fast Polynomial Transcendentals for LLMs

The paper investigates using short polynomial approximations to accelerate special‑function operations in large language models on NVIDIA Blackwell GPUs. By replacing native sigmoid, tanh, and SiLU with degree‑3 or degree‑4 bfloat16 programs, the authors achieve up to 2.19× speed‑ups in isolated FP16 benchmarks and modest training‑step throughput gains (2.7–8.0%) across four integration tasks. The study also evaluates model behavior, finding negligible training‑loss differences within 100 billion tokens.

By Robert Hu
arXiv AI
Jun 15

Gefen: Optimized Stochastic Optimizer

arXiv:2606. 13894v1 Announce Type: cross Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory.

By Nadav Benedek, Tomer Koren, Ohad Fried
arXiv AI
Jun 4

Stochastic Sparse Attention for Memory-Bound Inference

arXiv:2605. 01910v2 Announce Type: replace-cross Abstract: Autoregressive decoding becomes bandwidth-limited at long contexts, as generating each token requires reading all $n_k$ key and value vectors from KV cache.

By Kyle Lee, Corentin Delacour, Kevin Callahan-Coray, Kyle Jiang, Can Yaras, Samet Oymak, Tathagata Srimani, Kerem Y. Camsari
arXiv AI
Jul 29

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

arXiv:2607. 24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as matrix multiplication, convolution, and normalization.

By Joshua Brodsky, Dhravid Kumar, Savini Kashmira, Jayanaka Danatanarayana, Jason Mars, Krisztian Flautner, Lingjia Tang