Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
HiLRP introduces a unified attribution framework for Vision Transformers (ViTs) that addresses the challenges posed by diverse architectural designs. By decomposing ViT operations into four basic types—linear maps, bilinear mixing, normalization/gating, and reindexing—HiLRP applies conservation‑satisfying relevance rules, enabling reliable explanations across a wide range of backbones. The method outperforms 14 existing attribution techniques on 10 architectures, maintaining conservation and improving localization accuracy (0.97 Pointing) compared to competitors.
arXiv:2505. 15441v5 Announce Type: replace-cross Abstract: Natural images exhibit strong geometric regularities: local structures, such as edges, corners, and textures, appear in many orientations and mirror configurations.
arXiv:2609.23733v1 Announce Type: new Abstract: Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-v...
arXiv:2610.01785v1 Announce Type: cross Abstract: Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibit...
arXiv:2607. 02097v1 Announce Type: cross Abstract: Large kernel depthwise convolutions achieve strong performance but suffer from significant degradation as kernel size grows due to irregular memory access from gather-based computation; while Large Kernel Acceleration (LKA) helps on small feature maps, it becomes counterproductive on large feature maps, even slower than non-accelerated implementations.
FreeFlow is a hierarchical transformer for optical flow estimation that eliminates traditional flow-specific inductive biases such as correlation volumes, feature warping, and iterative refinement. It relies on a single feed-forward encoder–decoder architecture that integrates window attention for local processing, shifted-window attention for cross-window communication, and global attention at reduced resolution. This design allows the model to scale naturally with capacity, achieving state-of-the-art performance on Sintel, KITTI-2015, and Spring benchmarks while remaining memory efficient at 1080p inference.