arXiv Computer Vision By Ron Campos, Subhajit Maity, Xin Li, Srijan Das, Aritra Dutta

DnA: Denoising Attention for Visual Tasks

Read the original on arXiv Computer Vision →

The paper introduces Denoising Attention (DnA), a modification to the standard softmax activation used in multihead attention for visual perception tasks. DnA employs a positive query to highlight correct class features and a negative query to suppress irrelevant features, projecting these interactions into two distinct subspaces to enhance discriminability. Experiments with a ViT-B backbone show an absolute 0.8% improvement on ImageNet-1K and additional gains on video understanding and video LLM tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computation and Language
Aug 31

Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs

The paper introduces Semantic Head Specialization (SHS), a phenomenon where Vision Transformer (ViT) attention heads specialize as either object- or background-focused, most evident under full attention. It proposes the SHS-Index to quantify this specialization, demonstrating its ability to distinguish full-attention from chunk-window ViTs and its strong correlation with downstream benchmark performance. Leveraging insights into window interaction, token serialization, and local softmax allocation, the authors design Ariadne Attention, a hybrid attention mechanism that matches full-attention performance on 22 image and video tasks while reducing attention compute by 6.5×.

By Chenhong He, Lei Li, Shicheng Li, Hanglong Lv, Lingpeng Kong, Qi Liu, Tong Yang, Shuhuai Ren
Hugging Face Trending Papers
Sep 17

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Video DeltaNet (VDN) introduces a hybrid attention mechanism for video diffusion models, combining local Softmax attention with a bidirectional linear memory branch called Video Delta Attention (VDA). The design updates memory once per frame, uses separate output projections and learnable gates to balance the two branches, and employs a staged teacher‑alignment recipe to integrate the new pathway into pretrained models. When applied to MiniMax H3, VDN achieves a 14.5× speedup, completing 14.3‑second, 768p video denoising in 6.70 seconds on eight NVIDIA B200 GPUs compared to the 50‑step dense baseline.