ARROW: Arbitrary Reconstruction and Tracking of 4D Observations in the Wild
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires both geometric consistency and motion fidelity with respect to the reference video. Existing methods based on explicit 3D representations are limited by the accuracy of off-the-shelf reconstruction modules, which often produce inaccurate geometry for dynamic objects in monocular videos.
TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.
The paper introduces PLANET, a multi‑object tracker that transcends traditional image‑plane limitations by incorporating 3D scene geometry into its query formation. By lifting 2D tracking datasets into 3D and embedding reconstructed geometry into features and positional encodings, PLANET encourages queries to encode object positions. An auxiliary 3D location prediction task and a dual‑resolution temporal memory further enhance performance, enabling state‑of‑the‑art results on three diverse benchmarks.
TrackEverything is a 3D point tracker that overcomes the trade‑off between sparse long‑horizon tracking and dense short‑clip tracking by representing videos as persistent 3D scene tracks in world coordinates. It introduces voxel‑based de‑duplication at sliding‑window boundaries, a two‑stage refinement process (endpoint refiner and lightweight trajectory refiner), and a 3D WAFT module that replaces memory‑heavy 4D correlation volumes with efficient feature sampling. The method can track all visible points in videos longer than 1000 frames using only 40 GB of GPU memory, outperforming existing dense trackers on short clips and matching sparse trackers on long sequences.
Monocular videos record 3D scenes as sequences of 2D image-plane projections, obscuring depth and spatial relationships. Multi-object trackers localize and associate objects primarily using appearance...
arXiv:2609.18363v1 Announce Type: new Abstract: Online multi camera 3D tracking must maintain scene global identities across synchronized views, yet query-based trackers carry these identities only i...