arXiv:2509. 08421v2 Announce Type: replace-cross Abstract: For multimedia spatial intelligence through time, multi-view multi-object tracking (MVMOT) suffers from persistent challenges in maintaining consistent object identities across different camera views, leading to tracking inaccuracies.
By Keisuke Toida, Taigo Sakai, Takeshi Nakamura, Hiroshi Shimizu, Kazuhiro Hotta
arXiv:2608.20639v1 Announce Type: new
Abstract: Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view map from multi-view images. Recent MVPD methods adop...
By Taiga Yamane, Satoshi Suzuki, Ryo Masumura, Shota Orihashi, Tomohiro Tanaka, Mana Ihori, Naoki Makishima
arXiv:2609.18363v1 Announce Type: new
Abstract: Online multi camera 3D tracking must maintain scene global identities across synchronized views, yet query-based trackers carry these identities only i...
By Pragyan Shrestha, Haruto Nakayama, Atom Scott
The paper introduces SAM‑H, a planar object tracker that estimates 8‑degree‑of‑freedom homographies directly from segmentation mask contours using a training‑free pipeline. When applied to masks from SAM 2, SAM‑H achieves a new state‑of‑the‑art performance on the PlanarTrack benchmark, improving the p@5 metric by 18.4 percentage points. The authors also demonstrate that combining segmentation‑based and correspondence‑based homography estimation yields WOFTSAM, which surpasses all previous methods on both PlanarTrack and POT‑210, and provide precise re‑annotations of PlanarTrack initial poses for more accurate benchmarking.
By Jonas Serych, Jiri Matas
The paper introduces Post Fusion Stabilizer (PFS), a lightweight module that refines intermediate bird’s‑eye view (BEV) feature maps in existing camera‑LiDAR fusion detectors. PFS stabilizes feature statistics under domain shift, suppresses regions affected by sensor degradation, and adaptively restores weakened cues via residual correction, acting as a near‑identity transformation. On the nuScenes benchmark, PFS achieves state‑of‑the‑art robustness, notably improving camera dropout robustness by +1.2% and low‑light performance by +4.4% mAP while adding only 3.3 M parameters.
By Trung Tien Dong, Dev Thakkar, Arman Sargolzaei, Xiaomin Lin
TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.
By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow
The paper introduces an on-the-fly homography calibration system for multi-camera tracking that starts from coarse manual homographies and refines them using a centroid-based projection optimization (PO) on live detection metadata. PO continuously aligns ground-plane geometry without adding computational latency, enabling the system to adapt automatically to camera movements or environmental changes. The refined geometry feeds a bird's-eye-view tracker that fuses detections and unifies trajectories across zones while maintaining privacy safety and zero overhead.
By David Voihanski, Mor Sinai, Ben Zion Bobrovsky
arXiv:2606. 07708v1 Announce Type: cross Abstract: We introduce a dataset and benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections.
By Prakhar Bhardwaj, Simone Weikl, Kilian Mang, Elia Jonas Sandtner
arXiv:2606. 13509v1 Announce Type: cross Abstract: Indoor vision-based localization systems are affected by detection noise, occlusions, and limited camera coverage, leading to uncertainty at multiple stages of the pipeline.
By Mateo Toro Diz, Jonathan Hoss, Noah Klarmann
arXiv:2608.24544v1 Announce Type: new
Abstract: Many feature-based visual-inertial odometry (VIO) systems rely on sparse feature tracking, whose accuracy and robustness directly affect state estimati...
By Renbiao Jin, Danping Zou, Wenxian Yu
arXiv:2606. 07233v1 Announce Type: cross Abstract: LiDAR-based 3D Multi-Object Tracking (MOT) typically relies solely on geometric information, which is often insufficient to distinguish between targets during prolonged occlusions or in crowded human-populated environments.
By Eduardo Borges, Lu\'is Garrote, Urbano J. Nunes
The paper introduces a predictive tracking method for optical coherence tomography (OCT) that propagates positional updates across multiple landmarks to estimate a global 6D pose. By doing so, it achieves robust tracking at higher velocities, reporting root‑mean‑square errors below 1 mm for speeds up to 100 mm/s while tracking up to nine landmarks sequentially.
By Konrad Reuter, Suresh Guttikonda, Chaitali Uday Karekar, Christian Betz, Alexander Schlaefer