arXiv Computer Vision

GRACE: Geometry- and Ray-Aware Camera-Efficient Multi-View Pedestrian Tracking

GRACE is a camera‑efficient multi‑view pedestrian tracker that reduces the number of required cameras while maintaining high tracking accuracy. It combines volumetric‑guided fusion of homography‑based BEV features with 3D‑lifted features, uses ray conditioning to incorporate each camera’s viewing direction, and employs BEV Track Recovery to continue existing tracks with low‑confidence detections. On the WildTrack dataset, GRACE raises MOTA from 83.54 to 91.07 compared to the baseline TrackTacular.

arXiv AI
Aug 11

ATLASFusion: Aggregation Tracking with Location-Aware Sparse Fusion for Robust Spatio-Temporal Multi-View Pedestrian Tracking

arXiv:2509. 08421v2 Announce Type: replace-cross Abstract: For multimedia spatial intelligence through time, multi-view multi-object tracking (MVMOT) suffers from persistent challenges in maintaining consistent object identities across different camera views, leading to tracking inaccuracies.

By Keisuke Toida, Taigo Sakai, Takeshi Nakamura, Hiroshi Shimizu, Kazuhiro Hotta
arXiv Computer Vision
Aug 26

Segmentation-Guided Homography Estimation for Long-Term Planar Tracking

The paper introduces SAM‑H, a planar object tracker that estimates 8‑degree‑of‑freedom homographies directly from segmentation mask contours using a training‑free pipeline. When applied to masks from SAM 2, SAM‑H achieves a new state‑of‑the‑art performance on the PlanarTrack benchmark, improving the p@5 metric by 18.4 percentage points. The authors also demonstrate that combining segmentation‑based and correspondence‑based homography estimation yields WOFTSAM, which surpasses all previous methods on both PlanarTrack and POT‑210, and provide precise re‑annotations of PlanarTrack initial poses for more accurate benchmarking.

By Jonas Serych, Jiri Matas
arXiv AI
Sep 7

Post Fusion Bird's Eye View Feature Stabilization for Robust Multimodal 3D Detection

The paper introduces Post Fusion Stabilizer (PFS), a lightweight module that refines intermediate bird’s‑eye view (BEV) feature maps in existing camera‑LiDAR fusion detectors. PFS stabilizes feature statistics under domain shift, suppresses regions affected by sensor degradation, and adaptively restores weakened cues via residual correction, acting as a near‑identity transformation. On the nuScenes benchmark, PFS achieves state‑of‑the‑art robustness, notably improving camera dropout robustness by +1.2% and low‑light performance by +4.4% mAP while adding only 3.3 M parameters.

By Trung Tien Dong, Dev Thakkar, Arman Sargolzaei, Xiaomin Lin
arXiv Computer Vision
Sep 3

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.

By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow
arXiv AI
1d ago

On-the-Fly Homographies Calibration for Multi-Camera Tracking

The paper introduces an on-the-fly homography calibration system for multi-camera tracking that starts from coarse manual homographies and refines them using a centroid-based projection optimization (PO) on live detection metadata. PO continuously aligns ground-plane geometry without adding computational latency, enabling the system to adapt automatically to camera movements or environmental changes. The refined geometry feeds a bird's-eye-view tracker that fuses detections and unifies trajectories across zones while maintaining privacy safety and zero overhead.

By David Voihanski, Mor Sinai, Ben Zion Bobrovsky
arXiv Machine Learning
Jun 8

Does Appearance Help? A Systematic Study of Image-Based Re-Identification in Online 3D Multi-Pedestrian Tracking

arXiv:2606. 07233v1 Announce Type: cross Abstract: LiDAR-based 3D Multi-Object Tracking (MOT) typically relies solely on geometric information, which is often insufficient to distinguish between targets during prolonged occlusions or in crowded human-populated environments.

By Eduardo Borges, Lu\'is Garrote, Urbano J. Nunes
arXiv Computer Vision
Sep 11

Predictive Multi-Landmark OCT Tracking for Increased Motion Robustness

The paper introduces a predictive tracking method for optical coherence tomography (OCT) that propagates positional updates across multiple landmarks to estimate a global 6D pose. By doing so, it achieves robust tracking at higher velocities, reporting root‑mean‑square errors below 1 mm for speeds up to 100 mm/s while tracking up to nine landmarks sequentially.

By Konrad Reuter, Suresh Guttikonda, Chaitali Uday Karekar, Christian Betz, Alexander Schlaefer