arXiv Machine Learning

Component-Weighted Centroid Search for Exact Incremental BPE

arXiv Machine Learning
Sep 18

A Table-Free Index for Tapered Memoization Grids: Compact Out-of-Core Evaluation of Functions of Sorted Arguments

The paper presents a table‑free index for tapered memoization grids, enabling compact out‑of‑core evaluation of functions that depend on sorted arguments. By showing that the grid’s key set corresponds to multiset combinations, the authors derive a closed‑form O(d) ranking and unranking scheme that removes the need for large preprocessing tables and allows order‑free parallel construction. The resulting values‑only flat array uses significantly less memory than hash‑map memoization, offers faster query times once cache limits are exceeded, and remains operable with memory‑mapped storage beyond RAM.

By Tamal Maharaj
arXiv Machine Learning
Sep 3

Stream-CQSA: Exact Out-of-Memory Recovery for Attention

Stream-CQSA is an attention-level out‑of‑memory recovery framework that uses cyclic quorum set (CQS) decomposition to recursively split an infeasible attention call into independent subsequence tasks. Each task is executed with a compatible inner kernel and the local statistics are recomposed to recover the full attention output exactly, whether the wrapped kernel is exact or approximate. Compared with FlashAttention‑2, Stream‑CQSA achieves comparable 16‑bit forward‑output error and matches backward‑gradient error when FlashAttention‑2 fits in GPU memory, but it incurs higher runtime and continues to produce outputs beyond FlashAttention‑2’s sequence‑length boundary where FlashAttention‑2 OOMs. whyItMatters":"Stream‑CQSA turns memory‑capacity failures into recoverable executions, enabling large‑context language models to run beyond the limits of existing attention implementations without sacrificing correctness."

By Yiming Bian, Joshua M. Akey