arXiv Machine Learning By Zhibo Yang

Locality Does Not Imply Reachability: Boundary Repair in Block-Sparse Causal Attention

Read the original on arXiv Machine Learning →

arXiv:2606. 02680v1 Announce Type: new Abstract: Sparse causal attention is usually described by sequence locality: nearby tokens should remain easy to access, while distant tokens may be dropped to reduce cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 11

Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics

arXiv:2605. 04893v2 Announce Type: replace Abstract: When a language model processes a hallucinated response, its attention routing tends to fail in one of two shapes: over-concentrating on a narrow set of positions, or spreading so diffusely that relevance is diluted, and the shape of the failure carries diagnostic signal.

By Dominik Dahlem, Diego Maniloff, Mac Misiura
arXiv Computation and Language
Aug 28

TwinKV: A Composable Repair Pass for KV Cache Eviction via Pairwise Key Redundancy

TwinKV is a training‑free, attention‑free repair pass that identifies and swaps orphaned and redundant tokens in a KV cache, improving long‑context inference for small models. It works by detecting near‑duplicate keys and can be composed with existing eviction policies without altering their scoring rules. Experiments on Qwen3‑4B and Llama‑3.2‑1B across LongBench, LooGLE, RULER, and MMLU‑Pro show that TwinKV consistently improves performance for most configurations, especially at tighter compression ratios.

By Hong Chen, Yudong Zeng, Yongwei Huang, Zuhao Ouyang, Junyan Zhang, Xuming Hu
arXiv AI
Aug 24

BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers

BF1 is a deterministic block‑aligned dyadic sparse‑attention retrofit designed to reduce the cost of causal attention in long‑context transformers. It combines a small exact local neighborhood, a global first block, and logarithmically spaced historical blocks, achieving O(n log n) token interactions per layer with O(log n) communication depth. On an NVIDIA RTX PRO 6000 Blackwell GPU, BF1 outperforms dense attention for 2K–4K tokens and delivers up to a 10.91× prefill speedup at 32K tokens, while retrofitting eight of 28 Qwen3‑0.6B layers reduces first‑token latency by up to 15.3% at 32K tokens and yields the lowest perplexity among compared sparse and dense training protocols.

By Hina Dixit