arXiv AI

MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel

arXiv:2607. 19456v1 Announce Type: cross Abstract: We derive four memory-optimal inference artifacts for transformer attention using the Mathematics of Arrays (MoA), each following directly from the forward-pass Denotational Normal Form (DNF) of with the query-row index fixed to the current decode step.

arXiv Machine Learning
Sep 3

Stream-CQSA: Exact Out-of-Memory Recovery for Attention

Stream-CQSA is an attention-level out‑of‑memory recovery framework that uses cyclic quorum set (CQS) decomposition to recursively split an infeasible attention call into independent subsequence tasks. Each task is executed with a compatible inner kernel and the local statistics are recomposed to recover the full attention output exactly, whether the wrapped kernel is exact or approximate. Compared with FlashAttention‑2, Stream‑CQSA achieves comparable 16‑bit forward‑output error and matches backward‑gradient error when FlashAttention‑2 fits in GPU memory, but it incurs higher runtime and continues to produce outputs beyond FlashAttention‑2’s sequence‑length boundary where FlashAttention‑2 OOMs. whyItMatters":"Stream‑CQSA turns memory‑capacity failures into recoverable executions, enabling large‑context language models to run beyond the limits of existing attention implementations without sacrificing correctness."

By Yiming Bian, Joshua M. Akey
arXiv Computer Vision
2d ago

Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference

The paper introduces Right In-Place (RiP) convolution, a memory‑efficient strategy that corrects and generalizes previous in‑place convolution formulations to arbitrary stride, dilation, padding, and rectangular kernels. RiP aligns each layer’s input and output within a shared workspace, enabling safe, row‑major access with minimal memory overhead. Experiments on 10,000 random layers and 84 layers from 25 architectures show no corruption, matching or improving on existing herringbone workspaces while reducing memory usage by up to 24.8% and lowering peak activation memory on Raspberry Pi Pico MCUs by 12.5–33.3% without affecting cycle counts.

By Opegbemi Matthias Busoye, Tolulope Matthew Busoye, Eghonghon-aye Eigbe
arXiv AI
Jun 4

Stochastic Sparse Attention for Memory-Bound Inference

arXiv:2605. 01910v2 Announce Type: replace-cross Abstract: Autoregressive decoding becomes bandwidth-limited at long contexts, as generating each token requires reading all $n_k$ key and value vectors from KV cache.

By Kyle Lee, Corentin Delacour, Kevin Callahan-Coray, Kyle Jiang, Can Yaras, Samet Oymak, Tathagata Srimani, Kerem Y. Camsari
arXiv AI
Jun 19

StreamKL: Fast and Memory-Efficient KL Divergence for Boosting Attention Distillation

arXiv:2606. 20005v1 Announce Type: cross Abstract: Attention distillation, which trains one attention distribution to match another by minimizing their Kullback-Leibler (KL) divergence, is widely used in knowledge distillation, model compression, continual learning, and sparse-attention LLM training.

By Guangda Liu, Yiquan Wang, Chengwei Li, Wenhao Chen, Jing Lin, Yiwu Yao, Danning Ke, Wenchao Ding, Jieru Zhao
arXiv AI
Jun 12

MiniMax Sparse Attention

arXiv:2606. 13392v1 Announce Type: new Abstract: Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale.

By Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao
arXiv Machine Learning
5d ago

The KV Cache Is the New Memory Wall

The paper argues that for large‑context autoregressive language‑model inference, memory bandwidth—specifically the Key‑Value (KV) cache—becomes the limiting resource rather than arithmetic throughput. It analytically derives how arithmetic intensity decays with context length for NVIDIA H100, NVIDIA B200, and AMD MI300X, identifies crossover points where KV traffic overtakes weight traffic, and evaluates representative techniques across five compression domains. The study finds a three‑regime behavior: below the crossover, weight traffic dominates and KV compression offers little benefit; beyond it, KV traffic dominates and compression methods trade quality for bandwidth, with paging and prefix sharing being lossless but capacity‑limited, while quantization and eviction directly reduce bandwidth at the cost of accuracy. whyItMatters":"The work provides a unified analytical framework and a standardized protocol that enable consistent comparison of KV‑compression techniques across hardware and workloads, guiding practitioners in selecting appropriate methods for long‑context inference."

By Tejinder Singh
arXiv AI
Aug 18

FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy

arXiv:2608. 15602v1 Announce Type: cross Abstract: While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specialized hardware kernels, thus failing to unleash the full acceleration potential due to persistent reliance on expensive floating-point arithmetic or runtime dequantization overheads.

By Qingyao Yang, Runming Yang, He Xiao, Wendong Xu, Junyu Chen, Haobo Liu, Chenchen Ding, Ruihan Hu, Yik-Chung Wu, Ngai Wong