arXiv Machine Learning By Nicolaj Rux, Sebastian Neumayer

Fast Gauss Sums via Flash Attention

Read the original on arXiv Machine Learning →

The paper introduces a method to compute Gaussian kernel sums—central to many kernel-based machine learning techniques—using flash attention, a highly optimized softmax attention implementation. By applying two small input augmentations, the authors transform the normalized softmax reduction into an unnormalized Gauss sum, eliminating the need for custom GPU code. For feature dimensions greater than eight in fp16, this approach outperforms both compiled PyTorch code and PyKeOps kernels in speed, memory overhead, and accuracy, while maintaining linear memory scaling.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 22

FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax Attention

FlashBoB introduces an I/O‑efficient algorithm for exact backward‑over‑backward (BoB) in softmax attention, enabling precise second‑order differentiation without large intermediate tensors. By exploiting a hierarchical affine structure, the method confines computation to on‑chip tiles and limits off‑chip memory traffic, achieving θ(N² d²/M) HBM usage. Experiments show FlashBoB scales to sequence lengths of 262K on a single A100 GPU, outperforming prior exact baselines and FlashBack by up to 6.3×.

By Anthony Givans, Michael Crawshaw, Mingrui Liu
arXiv AI
Jun 19

StreamKL: Fast and Memory-Efficient KL Divergence for Boosting Attention Distillation

arXiv:2606. 20005v1 Announce Type: cross Abstract: Attention distillation, which trains one attention distribution to match another by minimizing their Kullback-Leibler (KL) divergence, is widely used in knowledge distillation, model compression, continual learning, and sparse-attention LLM training.

By Guangda Liu, Yiquan Wang, Chengwei Li, Wenhao Chen, Jing Lin, Yiwu Yao, Danning Ke, Wenchao Ding, Jieru Zhao
arXiv Machine Learning
1d ago

Fast Polynomial Transcendentals for LLMs

The paper investigates using short polynomial approximations to accelerate special‑function operations in large language models on NVIDIA Blackwell GPUs. By replacing native sigmoid, tanh, and SiLU with degree‑3 or degree‑4 bfloat16 programs, the authors achieve up to 2.19× speed‑ups in isolated FP16 benchmarks and modest training‑step throughput gains (2.7–8.0%) across four integration tasks. The study also evaluates model behavior, finding negligible training‑loss differences within 100 billion tokens.

By Robert Hu