We study relative position encoding for multi-view vision Transformers under camera heterogeneity, including varying fields of view (FoVs) or projection models. Existing rotary relative position encod...
arXiv:2601. 15275v3 Announce Type: replace-cross Abstract: We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-invariant attention with multi-frequency similarity, and can adapt to the geometry of the underlying 3D scene.
By Yu Wu, Minsik Jeon, Jen-Hao Rick Chang, Oncel Tuzel, Shubham Tulsiani
arXiv:2608.23206v1 Announce Type: new
Abstract: We study spherical occupancy profiles-the ray-wise occupancy probability profiles P(r) = T(r) o(r) distilled from multi-view 3D Gaussian reconstruction...
By YiHsuan Tsai
arXiv:2605.12938v2 Announce Type: replace-cross
Abstract: Video world models should predict future appearance in a way that remains consistent with 3D scene structure, camera motion, and lens geometr...
By Seonghyun Jin, Youngmin Kim, Sunwoo Park, Jong Chul Ye
arXiv:2607. 00417v1 Announce Type: cross Abstract: In the era of satellite constellations, multi-view optical satellite imagery is pivotal for Earth Observation (EO) and high-quality Digital Surface Model (DSM) reconstruction.
By Qiyan Luo, Yingdong Pi, Lekang Wen, Jie Yang, Xiaoyu Wang, Haiming Zhang, Mi Wang
The paper proposes a new sequence-to-sequence formulation for multi-view stereo (MVS) that jointly predicts 3D geometry for all input views using a global transformer architecture. It introduces ray‑map embeddings to inject camera parameters into image tokens and a unified global cost volume to capture 3D structure across all views. Experiments on public benchmarks demonstrate state‑of‑the‑art performance, outperforming both traditional MVS and feed‑forward reconstruction baselines.
By Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua