arXiv:2605. 26702v2 Announce Type: replace-cross Abstract: Reliable watermarking of panoramic imagery is fundamentally challenged by arbitrary 3D rotations.
By Pengzhen Chen, Yanwei Liu, Xiaoyan Gu, Antonios Argyriou, Wu Liu, Weiping Wang
arXiv:2606. 27864v1 Announce Type: cross Abstract: Vision transformers have become a dominant architecture for visual recognition.
By T\={\i}kun \^Ong, Georg B\"okman
arXiv:2606. 04108v1 Announce Type: cross Abstract: Single-view 3D generative models have achieved impressive visual quality, yet they are not designed to satisfy structural or functional requirements, and in practice, often fall short.
By Guangda Ji, Qimin Chen, Qinchan Li, Mingrui Zhao, Kai Wang, Hao Zhang
arXiv:2505. 21736v2 Announce Type: replace-cross Abstract: Translation equivariance is a central reason convolutional neural networks have been successful in computer vision.
By Siqi Fang, Zachary Schlamowitz, Andrew Bennecke, Daniel J. Tward
Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs.
Geometric foundation models, such as the Visual Geometry Grounded Transformer (VGGT), provide strong 3D priors from unposed images. However, such models operate purely in a feed-forward, deterministic regime, \ie~they cannot generate plausible geometry beyond what the input views directly support.