arXiv:2609.40244v1 Announce Type: new
Abstract: Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving effi...
By Yufei Wei, Shuhao Ye, Qi Wang, Xin Zheng, Qing Huang, Rong Xiong, Yue Wang
arXiv:2608.27529v1 Announce Type: new
Abstract: Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation...
By Jiarong Han, Jincheng Xiong, Yuzhou Liu, Linzhe Shi, Changjie Wu, Ning Guo, Mu Xu, Hang Zhang, Ming Qian
Anchor3R is a streaming 3D reconstruction framework that predicts window-relative poses and local geometry in the current‑frame coordinate system, forming a dense relative‑pose graph for online pose updates and loop‑aware motion averaging. It improves long‑horizon pose accuracy and dense reconstruction quality on indoor, outdoor, driving, and RGB‑D benchmarks, and generalizes from 48‑frame training sequences to streams exceeding 10,000 frames while keeping GPU memory bounded. The method addresses issues of train‑test mismatch, early‑anchor bias, and accumulated drift found in previous streaming models.
By Peilin Tao, Chong Cheng, Yuansen Du, Caiwei Song, Zhengqing Chen, Xiaoyang Guo, Wei Yin, Weiqiang Ren, Qian Zhang, Hainan Cui, Shuhan Shen
The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.
By Vladimir Yugay, Duy-Kien Nguyen, Theo Gevers, Cees G. M. Snoek, Martin R. Oswald
The paper introduces G2G, a method that leverages known intra-group geometry to estimate the relative 6-DoF pose between two image groups, a key problem in cross-sequence relocalization and multi-camera rig odometry. G2G keeps a frozen foundation model and adds three lightweight trainable modules—a perceiver resampler, a cross-group bridge with merged self-attention, and a multi-frame pose head—totaling about 32 M parameters, less than 6% of the full model. Evaluated on four diverse datasets covering indoor/outdoor simulation, real-world cross-season capture, and zero-shot sim-to-real transfer, G2G achieves state‑of‑the‑art accuracy on both pose estimation tasks while only requiring supervision from relative poses.
By Yufei Wei, Shuhao Ye, Chenxiao Hu, Yiyuan Pan, Dongyu Feng, Rong Xiong, Yue Wang, Yanmei Jiao
arXiv:2609.13733v1 Announce Type: new
Abstract: Stable and reliable 4D spatial understanding is fundamental for autonomous driving systems. While feedforward reconstruction networks can estimate came...
By Meng-Li Shih, Shih-Yang Su, Yuliang Zou, Hao Xiang, Haidong Zhu, Vincent Casser, Brian Curless, Dmitry Kalenichenko, Mingxing Tan, Dragomir Anguelov
The paper introduces a reliability-regulated trajectory optimization framework for progressive COLMAP‑free 3D Gaussian Splatting (3DGS). It uses a self‑supervised bidirectional cycle‑consistency mechanism to control camera trajectory estimation through forward motion propagation and retrospective trajectory correction, thereby reducing error compounding without external priors. Experiments on Tanks and Temples and CO3D‑V2 demonstrate improved camera trajectory accuracy and novel‑view rendering quality compared to existing unposed baselines.
By Zijian Wu, Jinliang Wang, Zidian Lin, Ying Song, Ziqian Lu, Hanjie Ma, Zhen Ye, Mingfeng Jiang
arXiv:2609.13504v1 Announce Type: new
Abstract: Recent developments in feed-forward 3D reconstruction resulted in models which can recover dense scene representations and camera motion solely from an...
By Tingjun Huang, Dmitry Rudshin, Mathieu Meyer, Pietro Bonazzi, Marc Pollefeys, Emilia Szyma\'nska
arXiv:2504.15776v2 Announce Type: replace
Abstract: Public autonomous driving datasets underpin the training and benchmarking of perception, mapping, and localization algorithms, yet residual inaccur...
By Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Rold\~ao, Dzmitry Tsishkou, Bingbing Liu, Cyrille Migniot, Pascal Vasseur, C\'edric Demonceaux
arXiv:2605.16981v3 Announce Type: replace
Abstract: Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TT...
By Kejun Ren, Lei Jin, Tianxin Huang, Lianming Xu, Li Wang
arXiv:2609.24482v1 Announce Type: new
Abstract: Monocular 3D human pose estimation (HPE) remains challenging due to depth ambiguity, occlu- sions, and the need for temporal consistency. While multi-v...
By Mena Kamel, Natalie Won, Amrut Sarangi, Sven Jager, Albert Pla Planas
arXiv:2604.28130v4 Announce Type: replace
Abstract: Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts join...
By Kehong Gong, Zhengyu Wen, Dao Thien Phong, Mingxi Xu, Weixia He, Qi Wang, Ning Zhang, Zhengyu Li, Guanli Hou, Dongze Lian, Xiaoyu He, Mingyuan Zhang, Hanwang Zhang