arXiv AI

BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers

BF1 is a deterministic block‑aligned dyadic sparse‑attention retrofit designed to reduce the cost of causal attention in long‑context transformers. It combines a small exact local neighborhood, a global first block, and logarithmically spaced historical blocks, achieving O(n log n) token interactions per layer with O(log n) communication depth. On an NVIDIA RTX PRO 6000 Blackwell GPU, BF1 outperforms dense attention for 2K–4K tokens and delivers up to a 10.91× prefill speedup at 32K tokens, while retrofitting eight of 28 Qwen3‑0.6B layers reduces first‑token latency by up to 15.3% at 32K tokens and yields the lowest perplexity among compared sparse and dense training protocols.

arXiv Machine Learning
Jul 20

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts

arXiv:2605. 15422v3 Announce Type: replace Abstract: Modern RL post-training methods such as GRPO and DAPO train on N response sequences of R tokens sampled from a shared prompt of P tokens, but standard FlashAttention replicates all P prompt tokens N times across both forward and backward passes -- duplicating compute and memory on identical hidden states.

By Jiading Gai, Shuai Zhang, Xiang Song, Bernie Wang, George Karypis
arXiv Machine Learning
Sep 18

Post-Boundary Bridge: Must Local Attention Go Global Between Global Layers?

The paper introduces Post-Boundary Bridge (PBB), a hybrid transformer architecture that keeps causal attention within blocks while adding direct connections across block boundaries. PBB focuses on within-block modeling and nearby exchange, delegating long-range communication to full-attention layers. Experiments on dense and mixture-of-experts models ranging from 205 million to 2.07 billion parameters show that PBB hybrids maintain near-full perplexity, competitive downstream performance, and improved source retrieval, while Flash-PBB achieves 1.82× faster decoding with half the local key‑value cache compared to Flash‑SWA.

By Zhibo Yang
arXiv Machine Learning
Jul 17

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

arXiv:2607. 14952v1 Announce Type: new Abstract: A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K tokens or below and rely on length generalization at deployment.

By Changhai Zhou, Kieran Liu, Yuhua Zhou, Qian Qiao, Jun Gao, Harry Zhang, Irvine Lu, Nolan Ho, Lucian Li, Andrew Lei, Cleon Cheng, Steven Chiang, Yihang Zeng, Di Zhang, Rio Yang, Kaijie Chen, Andrew Chen, Pony Ma, Weizhong Zhang, Cheng Jin
arXiv AI
Jun 12

MiniMax Sparse Attention

arXiv:2606. 13392v1 Announce Type: new Abstract: Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale.

By Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao
arXiv Machine Learning
Sep 21

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

Elastic Threshold Attention (ETA) is a trainable attention mechanism that dynamically predicts contextual thresholds from query representations, enabling selective pruning of KV cache tokens during long‑context decoding. By multiplicatively suppressing sub‑threshold logits during training, ETA avoids representation collapse and eliminates localized attention sinks, allowing a 1.45B model to match dense attention performance at roughly 85% training sparsity and 38% active decode density. At inference, a custom Triton kernel achieves up to 2.5× faster decoding on sequences up to 512K tokens, and an offline calibration step can further reduce compute by 27% by freezing per‑head thresholds.

By Themistoklis Haris, Henry Li, Maryam Karimzadehgan
arXiv Machine Learning
Sep 18

Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training

The paper introduces Block Parallelism (BP) and Context‑Sharded Block Parallelism (CSBP) to improve training efficiency for Block Diffusion Language Models (BDLMs) with long contexts. By assigning each corrupted‑block computation to a separate rank and sharding the shared clean sequence, CSBP reduces communication overhead and memory usage while preserving training semantics. Experiments on 16 H200 GPUs and 8 H100 GPUs show throughput gains of up to 1.61× and 7.59×, respectively, and higher benchmark pass rates in practical fine‑tuning scenarios.

By Tarun Suresh, Pranshu Chaturvedi, Hangoo Kang, Parth Shroff, Ishan S. Khare, Hermann Kumbong, Azalia Mirhoseini
arXiv Machine Learning
4d ago

Block Sparse Flash Attention

Block Sparse Flash Attention (BSFA) is a drop‑in replacement for FlashAttention that speeds up long‑context inference by pruning about 50% of computation and memory transfers. It selects the top‑k most important value blocks for each query using exact query‑key similarities and calibrated per‑layer, per‑head thresholds, requiring only a one‑time training‑free calibration. On Llama‑3.1‑8B, BSFA delivers up to 1.13× speedup on LongBench with a 1.1% accuracy drop and up to 1.24× on Needle‑in‑a‑Haystack retrieval with a 1% drop, while the attention kernel itself accelerates by up to 1.38×.

By Daniel Ohayon, Itay Lamprecht, Itay Hubara, Israel Cohen, Daniel Soudry, Noam Elata
arXiv Machine Learning
Sep 14

Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale

The paper investigates how block‑diffusion language models can use a constant‑size cache to enable efficient parallel decoding. By employing sequence mixers that summarize completed blocks into a reusable state and a block‑causal training objective, the authors pretrain three 3B block‑diffusion denoisers (attention, Mamba, and hybrid) on 300 B tokens. The resulting state‑space cache remains O(1) in memory and latency regardless of context length, yielding significant speed‑up and memory savings compared to traditional attention‑based caches, especially at very long sequences.

By Vaibhav Singh, Pierre-Andr\'e No\"el, Torsten Scholak, Eugene Belilovsky, Oleksiy Ostapenko