arXiv:2509. 08421v2 Announce Type: replace-cross Abstract: For multimedia spatial intelligence through time, multi-view multi-object tracking (MVMOT) suffers from persistent challenges in maintaining consistent object identities across different camera views, leading to tracking inaccuracies.
By Keisuke Toida, Taigo Sakai, Takeshi Nakamura, Hiroshi Shimizu, Kazuhiro Hotta
arXiv:2608.20639v1 Announce Type: new
Abstract: Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view map from multi-view images. Recent MVPD methods adop...
By Taiga Yamane, Satoshi Suzuki, Ryo Masumura, Shota Orihashi, Tomohiro Tanaka, Mana Ihori, Naoki Makishima
arXiv:2609.18363v1 Announce Type: new
Abstract: Online multi camera 3D tracking must maintain scene global identities across synchronized views, yet query-based trackers carry these identities only i...
By Pragyan Shrestha, Haruto Nakayama, Atom Scott
The paper introduces SAM‑H, a planar object tracker that estimates 8‑degree‑of‑freedom homographies directly from segmentation mask contours using a training‑free pipeline. When applied to masks from SAM 2, SAM‑H achieves a new state‑of‑the‑art performance on the PlanarTrack benchmark, improving the p@5 metric by 18.4 percentage points. The authors also demonstrate that combining segmentation‑based and correspondence‑based homography estimation yields WOFTSAM, which surpasses all previous methods on both PlanarTrack and POT‑210, and provide precise re‑annotations of PlanarTrack initial poses for more accurate benchmarking.
By Jonas Serych, Jiri Matas
The paper introduces Post Fusion Stabilizer (PFS), a lightweight module that refines intermediate bird’s‑eye view (BEV) feature maps in existing camera‑LiDAR fusion detectors. PFS stabilizes feature statistics under domain shift, suppresses regions affected by sensor degradation, and adaptively restores weakened cues via residual correction, acting as a near‑identity transformation. On the nuScenes benchmark, PFS achieves state‑of‑the‑art robustness, notably improving camera dropout robustness by +1.2% and low‑light performance by +4.4% mAP while adding only 3.3 M parameters.
By Trung Tien Dong, Dev Thakkar, Arman Sargolzaei, Xiaomin Lin
TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.
By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow