arXiv:2608.22039v1 Announce Type: new
Abstract: Structure-from-Motion (SfM) is a cornerstone of 3D perception, yet current methods often fail when applied to complex videos involving challenging came...
By Sara Sabour, Linyi Jin, Richard Tucker, Amir Hertz, Marcus Brubaker, Saurabh Saxena, Junhwa Hur, Andrea Tagliasacchi, Deqing Sun, David J. Fleet, Richard Szeliski, Noah Snavely
The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.
By Vladimir Yugay, Duy-Kien Nguyen, Theo Gevers, Cees G. M. Snoek, Martin R. Oswald
arXiv:2609.40244v1 Announce Type: new
Abstract: Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving effi...
By Yufei Wei, Shuhao Ye, Qi Wang, Xin Zheng, Qing Huang, Rong Xiong, Yue Wang
arXiv:2607. 02360v3 Announce Type: replace-cross Abstract: Monocular spacecraft 6D pose estimation remains difficult under weak texture, thin structures, illumination variation, and occlusion.
By Zongwu Xie, Yonglong Zhang, Yifan Yang, Yang Liu, Guanghu Xie
VideoReloc presents a method for long‑term indoor video relocalization that relies on a compact semantic scene graph rather than visual appearance. By adaptively selecting clip lengths based on odometry and object‑motion criteria, the system gathers spatial evidence, verifies poses through object triplets, and refines orientation using box faces and gravity cues. This approach achieves high localization accuracy with a tiny 100 kB map, outperforming traditional appearance‑based methods on RIO10 and ReplicaCAD datasets.
By Qianru Li, Xuyang Chen, Xuqin Wang, Zhenghao Zhang, Hongyi Luo, Tao Wu, Daniel Cremers, Lu Liu, Yanfeng Zhang
Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present...
DAPEVO is a learned visual odometry system that independently estimates image and event correspondences at shared patch locations and fuses their correlation evidence before motion refinement. It maintains image and event descriptors for each tracked patch, using a learned scalar gate to combine modality-specific correlation embeddings for each patch–frame edge, followed by a shared recurrent refinement and bundle‑adjustment update. The method supports event‑only observations and modality‑aware keyframe culling, achieving low trajectory error even when RGB frames are sparse or degraded, outperforming DPVO, RAMP‑VO, and event‑only DEVO on UZH‑FPV and TartanEvent datasets.
By Luca Gandolfi, Simone Nascivera, Roberto Pellerito, Rong Zou, Chiara Plizzari, Davide Scaramuzza
arXiv:2603.25175v2 Announce Type: replace
Abstract: Monocular egocentric 3D pose estimation is difficult because severe foreshortening, self-occlusion, and a restricted field of view often remove the...
By Md Mushfiqur Azam, John Quarles, Kevin Desai
arXiv:2604.28130v4 Announce Type: replace
Abstract: Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts join...
By Kehong Gong, Zhengyu Wen, Dao Thien Phong, Mingxi Xu, Weixia He, Qi Wang, Ning Zhang, Zhengyu Li, Guanli Hou, Dongze Lian, Xiaoyu He, Mingyuan Zhang, Hanwang Zhang
The paper introduces EgoGenEval, a new benchmark that assesses the physical consistency of visual generators under ego‑motion by measuring Camera Motion Grounding and Scene State Preservation across 1,400 cases and 2,360 target views. Experiments on 16 pose‑free generators and two pose‑conditioned references show that current models struggle to maintain both camera motion and scene state simultaneously. A follow‑up study using EgoGen‑Train demonstrates that pairwise supervision does not effectively improve both metrics together, suggesting the need for a trajectory‑centric training paradigm.
By Yilin Long, Chenming Zhu, Zitang Gou, Jingli Lin, Tai Wang
arXiv:2509. 08421v2 Announce Type: replace-cross Abstract: For multimedia spatial intelligence through time, multi-view multi-object tracking (MVMOT) suffers from persistent challenges in maintaining consistent object identities across different camera views, leading to tracking inaccuracies.
By Keisuke Toida, Taigo Sakai, Takeshi Nakamura, Hiroshi Shimizu, Kazuhiro Hotta
arXiv:2609.38755v1 Announce Type: new
Abstract: A wide range of approaches have been developed for camera pose estimation, including correspondence-based methods, end-to-end pose regression, and rece...
By Zhining Gu, Shangjie Du, Weimin Qiu, Carl Olsson, Ping Liu, Meng Tang