Rethinking the State Update Gate for Long-Sequence Recurrent 3D Reconstruction
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608.27529v1 Announce Type: new Abstract: Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation...
Streaming 3D reconstruction relies on a compact recurrent scene state to process long image streams in linear time and bounded memory. However, repeated updates can gradually corrupt this state, causing reliable historical information to be overwritten by noisy or ambiguous observations.
Anchor3R is a streaming 3D reconstruction framework that predicts window-relative poses and local geometry in the current‑frame coordinate system, forming a dense relative‑pose graph for online pose updates and loop‑aware motion averaging. It improves long‑horizon pose accuracy and dense reconstruction quality on indoor, outdoor, driving, and RGB‑D benchmarks, and generalizes from 48‑frame training sequences to streams exceeding 10,000 frames while keeping GPU memory bounded. The method addresses issues of train‑test mismatch, early‑anchor bias, and accumulated drift found in previous streaming models.
LoG-VGGT is a memory‑efficient framework for long‑sequence 3D reconstruction that balances local temporal modeling with global camera consistency. It uses cross‑window attention in a small subset of transformer blocks to propagate information across adjacent temporal windows while keeping memory usage bounded. A global camera consistency refinement module further improves long‑horizon pose stability by enforcing scene‑level constraints through cross‑attention between camera and compact register tokens, leading to better depth accuracy and robust camera pose estimation on multiple benchmarks.
Feed-forward 3D Gaussian Splatting enables efficient novel-view synthesis without per-scene optimization, but most existing methods assume a fixed set of context views and process them jointly. This limits their applicability to online scenarios where calibrated views arrive sequentially and the scene must be updated causally.
arXiv:2603.01765v5 Announce Type: replace Abstract: Monocular depth foundation models generalize across diverse scenes, but recovering accurate metric depth consistent with a target sensor remains ch...