StreamRig: Exploiting Intra-Rig Geometry for Streaming Multi-Camera Odometry
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present...
arXiv:2608.27529v1 Announce Type: new Abstract: Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation...
The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.
Anchor3R is a streaming 3D reconstruction framework that predicts window-relative poses and local geometry in the current‑frame coordinate system, forming a dense relative‑pose graph for online pose updates and loop‑aware motion averaging. It improves long‑horizon pose accuracy and dense reconstruction quality on indoor, outdoor, driving, and RGB‑D benchmarks, and generalizes from 48‑frame training sequences to streams exceeding 10,000 frames while keeping GPU memory bounded. The method addresses issues of train‑test mismatch, early‑anchor bias, and accumulated drift found in previous streaming models.
The paper introduces G2G, a method that leverages known intra-group geometry to estimate the relative 6-DoF pose between two image groups, a key problem in cross-sequence relocalization and multi-camera rig odometry. G2G keeps a frozen foundation model and adds three lightweight trainable modules—a perceiver resampler, a cross-group bridge with merged self-attention, and a multi-frame pose head—totaling about 32 M parameters, less than 6% of the full model. Evaluated on four diverse datasets covering indoor/outdoor simulation, real-world cross-season capture, and zero-shot sim-to-real transfer, G2G achieves state‑of‑the‑art accuracy on both pose estimation tasks while only requiring supervision from relative poses.
The paper introduces a reliability-regulated trajectory optimization framework for progressive COLMAP‑free 3D Gaussian Splatting (3DGS). It uses a self‑supervised bidirectional cycle‑consistency mechanism to control camera trajectory estimation through forward motion propagation and retrospective trajectory correction, thereby reducing error compounding without external priors. Experiments on Tanks and Temples and CO3D‑V2 demonstrate improved camera trajectory accuracy and novel‑view rendering quality compared to existing unposed baselines.