arXiv Machine Learning By Yiming Bian, Joshua M. Akey

Stream-CQSA: Exact Out-of-Memory Recovery for Attention

Read the original on arXiv Machine Learning →

Stream-CQSA is an attention-level out‑of‑memory recovery framework that uses cyclic quorum set (CQS) decomposition to recursively split an infeasible attention call into independent subsequence tasks. Each task is executed with a compatible inner kernel and the local statistics are recomposed to recover the full attention output exactly, whether the wrapped kernel is exact or approximate. Compared with FlashAttention‑2, Stream‑CQSA achieves comparable 16‑bit forward‑output error and matches backward‑gradient error when FlashAttention‑2 fits in GPU memory, but it incurs higher runtime and continues to produce outputs beyond FlashAttention‑2’s sequence‑length boundary where FlashAttention‑2 OOMs. whyItMatters":"Stream‑CQSA turns memory‑capacity failures into recoverable executions, enabling large‑context language models to run beyond the limits of existing attention implementations without sacrificing correctness."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 4

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

arXiv:2608. 01651v1 Announce Type: cross Abstract: Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound.

By Li Wang, Yi Su, Xiabao Wu, Chiran You, Yongchao Liu, Zhan Qiu, Juelu Zhang, Jiajun Zheng, Fangxin Liu, Jie Zhang, Chen Tian, Chengying Huan