arXiv:2606. 30576v1 Announce Type: cross Abstract: Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.
By Liyao Wang, Ruipu Wu, Haojun Xu, Lei Shi, Linjiang Huang, Si Liu
TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.
By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow
FreeInpaint is a feed‑forward 3D inpainting framework that reconstructs complete, geometrically consistent scenes directly from unposed multi‑view images with masked regions. It extends a 3D foundation model by propagating masked areas across views, using a Learnable Mask Attention mechanism to maintain reliable cross‑view correspondences and a Support Token Refinement strategy that injects diffusion‑generated auxiliary tokens for high‑fidelity completion. Experiments on diverse datasets show that FreeInpaint delivers superior inpainting quality without requiring pre‑computed camera poses, while maintaining fast inference speed.
By Jingyi Pan, Dan Xu, Qiong Luo
arXiv:2603.12064v3 Announce Type: replace
Abstract: We address the challenging problem of dense dynamic scene reconstruction and camera pose estimation from multiple freely moving cameras -- a settin...
By Shuo Sun, Unal Artan, Malcolm Mielle, Achim J. Lilienthaland, Martin Magnusson
SymRegFlow is a symmetry‑regularized flow‑matching framework that enables multi‑view‑consistent video generation across continuously varying camera poses without requiring ground‑truth novel‑view RGB supervision. The method geometrically warps source views into noisy anchors and uses masked dual‑anchor supervision combined with cross‑anchor denoising‑output consistency to reduce anchor‑specific errors. Experiments on Cosmos‑Drive‑Dreams and nuScenes show that SymRegFlow achieves superior video quality, achieving the lowest FVD and FVMD scores and improving FID and instance preservation compared to existing baselines.
By Xi Ye, Yuzhu Wang, Xiaoyang Liu, Jiayi Wang, Yangyang Xu, Ruyu Wang, Wenlin Chen, Duo Su, Jun Zhu
arXiv:2609.39116v1 Announce Type: new
Abstract: Prior-free 6D object pose tracking seeks to recover the trajectory of an unseen object from a single RGB video without object-specific CAD models, pose...
By Shiyang Liu, Weiquan Lin, Luping Xiao, Jiadong Tang, Yi Yang, Yu Gao, Xingyu Chen