arXiv:2606. 25318v1 Announce Type: cross Abstract: In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention.
By Sheir A. Zaheer, Alexander C. Holston, Chan Y. Park
arXiv:2606. 27864v1 Announce Type: cross Abstract: Vision transformers have become a dominant architecture for visual recognition.
By T\={\i}kun \^Ong, Georg B\"okman
arXiv:2505. 21736v2 Announce Type: replace-cross Abstract: Translation equivariance is a central reason convolutional neural networks have been successful in computer vision.
By Siqi Fang, Zachary Schlamowitz, Andrew Bennecke, Daniel J. Tward
arXiv:2607. 04262v1 Announce Type: new Abstract: Convolutional Neural Network (CNN) and Vision Transformer (ViT) for image classification exploit a dense grid of pixels containing redundant information.
By Sarabeshwar Balaji, Shubham Mohanty, Akash Anil
Vision Transformers underperform convolutional networks when training data is scarce, and distilling convolutional inductive biases from a CNN teacher is an effective remedy that leaves the deployed model unchanged. General-purpose feature distillation, however, transfers little in this setting.
arXiv:2509. 11218v2 Announce Type: replace-cross Abstract: Spatial transformations such as rotation and scale obscure the morphological cues needed for accurate image classification.
By Johann Schmidt, Sebastian Stober