We study relative position encoding for multi-view vision Transformers under camera heterogeneity, including varying fields of view (FoVs) or projection models. Existing rotary relative position encod...
The paper introduces G-ray, a ray-level relative position encoding for multi-view vision Transformers that remains consistent across different camera projections. By parameterizing rotary phases with camera-local ray angles, G-ray achieves projection-invariant positional consistency and can be integrated with existing encodings without extra learned parameters. Experiments on 3D reconstruction and novel-view synthesis benchmarks show that G-ray improves performance, notably reducing mean pointmap relative error by 45.8% over MapAnything.
By Shuo Zhang, Xin Su, Wei Wang, Jun Liu, Xinrui Zeng, Yongsen Chen, Chenjie Wang, Guibo Zhu, Jinqiao Wang, Bin Luo, Liangpei Zhang
arXiv:2606. 17564v1 Announce Type: cross Abstract: Standardized evaluation protocols are indispensable for robust benchmarking in remote sensing, particularly as foundation features are increasingly transferred across diverse sensors and complex imaging geometries.
By Qiyan Luo, Jie Yang, Yingdong Pi, Lekang Wen, Mi Wang
GeoCond is a lightweight reliability adapter that enhances frozen feed‑forward 3D reconstruction backbones by reading their predicted geometry to produce pose‑level uncertainty and a refinement gate. It can be trained using permutation‑orbit variance, ground‑truth pose error, or cycle residuals from unlabeled pose graphs, and at inference requires only a single backbone pass plus a small MLP. On the VGGT backbone, GeoCond reduces out‑of‑distribution AUSE from 0.32 to 0.20, transfers zero‑shot to outdoor extreme‑view scenes, and prevents collapse from uniform bundle adjustment, while also enabling gated refinement, pose‑graph weighting, calibration, curation, and capture decisions.
By David Ahmedt-Aristizabal, Mohammad Ali Armin, Russell Tsuchida, Lars Petersson
Feed-forward 3D foundation models such as VGGT predict cameras, depth, and point maps in a single pass, but can fail silently under low overlap, low parallax, and extreme relative rotation. Stratified...
arXiv:2601. 15275v3 Announce Type: replace-cross Abstract: We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-invariant attention with multi-frequency similarity, and can adapt to the geometry of the underlying 3D scene.
By Yu Wu, Minsik Jeon, Jen-Hao Rick Chang, Oncel Tuzel, Shubham Tulsiani
arXiv:2608.29680v1 Announce Type: new
Abstract: Feed-forward 3D foundation models reconstruct perspective scenes in one pass. Satellite photogrammetry needs a different product, one that domain adapt...
By Zhe Dong, Wanqing Wu, Yuzhe Sun, Haochen Jiang, Yuchen Ma, Lecheng Ren, Tianzhu Liu, Yanfeng Gu
ARC‑Loc introduces a new cross‑view localization method that bypasses heavy Bird’s‑Eye‑View transformations and external depth models. By converting ground keypoints into azimuthal rays on a satellite map and exploiting their convergence at the user’s location, the approach uses a minimal Azimuthal Ray Convergence solver and an ARC loss to directly match ground and satellite images. Experiments on VIGOR and KITTI show that ARC‑Loc achieves competitive accuracy while offering faster, memory‑efficient inference and easy integration with existing frameworks.
By Hyeongsik Kim, Mincheol Kim, Heejoon Moon, Je Hyeong Hong
arXiv:2605.12938v2 Announce Type: replace-cross
Abstract: Video world models should predict future appearance in a way that remains consistent with 3D scene structure, camera motion, and lens geometr...
By Seonghyun Jin, Youngmin Kim, Sunwoo Park, Jong Chul Ye
arXiv:2606. 24799v1 Announce Type: cross Abstract: Generic text-to-video models can be used as rich open-world scene priors.
By Chenrui Fan, Paolo Favaro
The paper introduces GRF-Recon, a framework for stable and scalable feed-forward 3D reconstruction from long monocular image sequences. It combines coarse-to-fine trajectory alignment, lightweight geometric prior injection via LoRA adaptation, and a hybrid-weight sparse ray-field optimization to refine local point clouds while enforcing cross-frame consistency. An efficient trajectory stitching strategy with joint ray-error optimization further reduces accumulated drift, achieving competitive trajectory accuracy compared to SLAM systems while maintaining globally consistent reconstructions in large-scale scenarios.
By Enpeng Li, Yunzhou Zhang, Zhiyao Zhang, Dexuan Lyu, Chenyu Wang, Chiyuan Cui, Cheng Cheng
arXiv:2608.23206v1 Announce Type: new
Abstract: We study spherical occupancy profiles-the ray-wise occupancy probability profiles P(r) = T(r) o(r) distilled from multi-view 3D Gaussian reconstruction...
By YiHsuan Tsai