arXiv AI By Lenore Mulin, Gaetan Hains

MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel

Read the original on arXiv AI →

arXiv:2607. 19456v1 Announce Type: cross Abstract: We derive four memory-optimal inference artifacts for transformer attention using the Mathematics of Arrays (MoA), each following directly from the forward-pass Denotational Normal Form (DNF) of with the query-row index fixed to the current decode step.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 3

Stream-CQSA: Exact Out-of-Memory Recovery for Attention

Stream-CQSA is an attention-level out‑of‑memory recovery framework that uses cyclic quorum set (CQS) decomposition to recursively split an infeasible attention call into independent subsequence tasks. Each task is executed with a compatible inner kernel and the local statistics are recomposed to recover the full attention output exactly, whether the wrapped kernel is exact or approximate. Compared with FlashAttention‑2, Stream‑CQSA achieves comparable 16‑bit forward‑output error and matches backward‑gradient error when FlashAttention‑2 fits in GPU memory, but it incurs higher runtime and continues to produce outputs beyond FlashAttention‑2’s sequence‑length boundary where FlashAttention‑2 OOMs. whyItMatters":"Stream‑CQSA turns memory‑capacity failures into recoverable executions, enabling large‑context language models to run beyond the limits of existing attention implementations without sacrificing correctness."

By Yiming Bian, Joshua M. Akey
arXiv Computer Vision
2d ago

Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference

The paper introduces Right In-Place (RiP) convolution, a memory‑efficient strategy that corrects and generalizes previous in‑place convolution formulations to arbitrary stride, dilation, padding, and rectangular kernels. RiP aligns each layer’s input and output within a shared workspace, enabling safe, row‑major access with minimal memory overhead. Experiments on 10,000 random layers and 84 layers from 25 architectures show no corruption, matching or improving on existing herringbone workspaces while reducing memory usage by up to 24.8% and lowering peak activation memory on Raspberry Pi Pico MCUs by 12.5–33.3% without affecting cycle counts.

By Opegbemi Matthias Busoye, Tolulope Matthew Busoye, Eghonghon-aye Eigbe