Profiling in PyTorch (Part 3): Attention is all you profile
Related stories
Profiling in PyTorch (Part 2): From nn.Linear to a Fused MLP
Visualize and understand GPU memory in PyTorch
OpenAI standardizes on PyTorch
We are standardizing OpenAI’s deep learning framework on PyTorch.
nanoVLM: The simplest repository to train your VLM in pure PyTorch
CBAM Paper Walkthrough: The Double-Attention Mechanism
Understanding and implementing CBAM (Convolutional Block Attention Module) from scratch with PyTorch The post CBAM Paper Walkthrough: The Double-Attention Mechanism appeared first on Towards Data Scie...
Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs
The paper introduces Semantic Head Specialization (SHS), a phenomenon where Vision Transformer (ViT) attention heads specialize as either object- or background-focused, most evident under full attention. It proposes the SHS-Index to quantify this specialization, demonstrating its ability to distinguish full-attention from chunk-window ViTs and its strong correlation with downstream benchmark performance. Leveraging insights into window interaction, token serialization, and local softmax allocation, the authors design Ariadne Attention, a hybrid attention mechanism that matches full-attention performance on 22 image and video tasks while reducing attention compute by 6.5×.
Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?
The paper investigates whether recent attention‑mechanism improvements—specifically gated attention, Kimi K3, Kimi Delta Attention, and Attention Residuals—effectively eliminate the attention‑sink problem when scaling language models to a one‑million‑token context window. Using a new diagnostic suite called SinkProbe, the authors evaluate sink mass, massive activation, position‑resolved recall, and the recency gap across four small models that vary only in token mixing and depth. Their findings show that the training objective, rather than the architecture, drives the emergence of attention sinks; gating did not replicate its previously reported benefits at the larger scale, and sink mass, activations, and positional bias behaved independently.
Accelerating PyTorch distributed fine-tuning with Intel technologies
Block Sparse Flash Attention
Block Sparse Flash Attention (BSFA) is a drop‑in replacement for FlashAttention that speeds up long‑context inference by pruning about 50% of computation and memory transfers. It selects the top‑k most important value blocks for each query using exact query‑key similarities and calibrated per‑layer, per‑head thresholds, requiring only a one‑time training‑free calibration. On Llama‑3.1‑8B, BSFA delivers up to 1.13× speedup on LongBench with a 1.1% accuracy drop and up to 1.24× on Needle‑in‑a‑Haystack retrieval with a 1% drop, while the attention kernel itself accelerates by up to 1.38×.
Limitations of Normalization in Attention Mechanism
arXiv:2508. 17821v3 Announce Type: replace-cross Abstract: This paper investigates the limitations of the normalization in attention mechanisms.
Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing
The paper investigates how the topology of attention graphs can differentiate hallucinated from non-hallucinated responses in large language models. By analyzing Forman-Ricci curvature, the authors identify structural bottlenecks and develop a method that captures both semi-local and global information-flow characteristics of attention heads associated with hallucinations. Extensive evaluation across multiple LLMs and benchmarks shows that this single-pass approach consistently outperforms existing attention-based and multi-response baselines, while also revealing that impaired context sharing—such as over-reliance on self-attention and information over-squashing—correlates strongly with hallucination occurrences.