In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant networks preserve the rotational, flip and positional symmetry in feature maps, making them useful for tasks where orientation of the inputs is relevant to the model outputs.
arXiv:2608.30417v1 Announce Type: new
Abstract: We give a complete characterization of equivariant multi-head self-attention (MHSA): if an MHSA layer is equivariant to a symmetry group $G$, then $G$...
By T\=ikun \^Ong
arXiv:2505. 15441v5 Announce Type: replace-cross Abstract: Natural images exhibit strong geometric regularities: local structures, such as edges, corners, and textures, appear in many orientations and mirror configurations.
By David Nordstr\"om, Johan Edstedt, Fredrik Kahl, Georg B\"okman
arXiv:2606. 25318v1 Announce Type: cross Abstract: In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention.
By Sheir A. Zaheer, Alexander C. Holston, Chan Y. Park
arXiv:2607. 15536v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) captures scenes by coupling explicit geometry (position, covariance) with view-dependent photometry (Spherical Harmonics).
By Chankyo Kim, Maani Ghaffari
arXiv:2609.01129v1 Announce Type: new
Abstract: We identify a recurrent algebraic regularity in Transformer attention: a sparse subset of effective OV operators $T=OV^\top$ nearly closes under compos...
By Jiming Feng, Junliang Li
arXiv:2607. 00556v1 Announce Type: cross Abstract: While recent advancements like the Poincar\'e ResNet have demonstrated the potential of learning visual representations directly in hyperbolic space, their optimisation remains hampered by the computationally intensive nature of Riemannian gradients and the strict boundaries of the manifold.
By Aiden Durrant, Rahul Baburajan, Georgios Leontidis
arXiv:2610.00848v1 Announce Type: cross
Abstract: Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoreg...
By Shao-Jun Xia, Huixin Zhang, Zhen Lei, Anlan Sun, Yuner Zhang, Xiaoyang Chen
arXiv:2607. 02386v1 Announce Type: cross Abstract: While Vision Transformers have achieved remarkable success across computer vision and language applications, the geometric evolution of their internal representations throughout training remains insufficiently understood.
By Kaustubh Kapil, Kishor P. Upla
arXiv:2505. 21736v2 Announce Type: replace-cross Abstract: Translation equivariance is a central reason convolutional neural networks have been successful in computer vision.
By Siqi Fang, Zachary Schlamowitz, Andrew Bennecke, Daniel J. Tward
arXiv:2609.23733v1 Announce Type: new
Abstract: Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-v...
By Abteen Arab, Guile Wu, Chengjie Huang, Dongfeng Bai
arXiv:2207. 03116v4 Announce Type: replace Abstract: We introduce a general method for learning representations that are equivariant to symmetries of data.
By Giovanni Luca Marchetti, Gustaf Tegn\'er, Anastasiia Varava, Danica Kragic