arXiv:2607. 04262v1 Announce Type: new Abstract: Convolutional Neural Network (CNN) and Vision Transformer (ViT) for image classification exploit a dense grid of pixels containing redundant information.
By Sarabeshwar Balaji, Shubham Mohanty, Akash Anil
Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs.
arXiv:2505. 21736v2 Announce Type: replace-cross Abstract: Translation equivariance is a central reason convolutional neural networks have been successful in computer vision.
By Siqi Fang, Zachary Schlamowitz, Andrew Bennecke, Daniel J. Tward
arXiv:2510. 03511v3 Announce Type: replace-cross Abstract: While widespread, Transformers lack inductive biases for geometric symmetries common in science and computer vision.
By Mohammad Mohaiminul Islam, Rishabh Anand, David R. Wessels, Friso de Kruiff, Thijs P. Kuipers, Rex Ying, Clara I. S\'anchez, Sharvaree Vadgama, Georg B\"okman, Erik J. Bekkers
arXiv:2606. 14597v1 Announce Type: new Abstract: Transformer-based neural operators have shown remarkable performance for approximating solution operators of partial differential equations on complex geometries.
By Armand de Villeroch\'e, Sibo Cheng, Vincent Le Guen, Marc Bocquet, Rem-Sophia Mouradi, Patrick Armand, Alban Farchi, Patrick Massin
Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields. Despite their remarkable capacity for representing geometric structures, ENNs suffer from degraded expressivity when processing symmetric inputs: the output representations are invariant to transformations that extend beyond the input's symmetries.
arXiv:2606. 24178v1 Announce Type: cross Abstract: Pretrained vision models often misclassify inputs that are rotated, scaled, or sheared, even though these affine transformations leave the object class unchanged.
By Dominik Lindner, Johann Schmidt, Tom Siegl, Martin Becker, Sebastian Stober
arXiv:2606. 17961v1 Announce Type: cross Abstract: Positional encoding is a fundamental component of Transformer architectures, as it injects information about the spatial or sequential arrangement of inputs.
By Andrea Santomauro, Luigi Portinale, Giorgio Leonardi
arXiv:2505. 19809v3 Announce Type: replace-cross Abstract: In many real-world applications of regression, conditional probability estimation, and uncertainty quantification, exploiting symmetries rooted in physics or geometry can dramatically improve generalization and sample efficiency.
By Daniel Ordo\~nez-Apraez, Vladimir Kosti\'c, Alek Fr\"ohlich, Vivien Brandt, Karim Lounici, Massimiliano Pontil
arXiv:2608. 12010v1 Announce Type: new Abstract: Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields.
By Ning Lin, Jiacheng Cen, Anyi Li, Wenbing Huang, Hao Sun
arXiv:2606. 29579v1 Announce Type: cross Abstract: Spatial reasoning remains a persistent challenge for many vision language models (VLMs), and improving it typically requires fine-tuning with substantial additional parameters.
By Rahul Chowdhury, Timothy A Rupprecht, Xuan Shen, Pu Zhao, Yanzhi Wang
arXiv:2312. 08230v2 Announce Type: replace-cross Abstract: Detecting partial extrinsic symmetry in 3D geometry is a fundamental yet persistent challenge in computer vision and graphics, critical for tasks ranging from shape completion to procedural generation.
By Gregor Kobsik, Isaak Lim, Leif Kobbelt