ReWorld-Track introduces a recursive event world model for language‑guided multi‑camera tracking that explicitly carries association uncertainty into future predictions. By treating candidate matches and waiting as alternative target states, the model updates a persistent recurrent belief that preserves uncertainty across successive observations. This approach improves identity continuity and next‑camera accuracy, achieving HOTA scores of 65.19 on CityFlowV2 and 45.36 on MTMMC, and reducing median arrival‑time error from 0.78 s to 0.71 s.
By Haoyang Wu, Shoudong Han, Chaoyue Li, Sijia Chen, Zhenyang Xie, Wang sihan
arXiv:2609.12261v1 Announce Type: new
Abstract: Multi-object tracking (MOT) is dominated by the tracking-by-detection paradigm, whose methods typically rely on a small set of hyperparameters that are...
By Momir Ad\v{z}emovi\'c
The paper introduces Privileged Appearance Transfer for Tracking (PATT), a teacher‑student framework that leverages exact target crops from past, current, and future frames during training to improve visual tracking. By weighting the teacher’s guidance with its localization advantage and accuracy, PATT transfers privileged appearance information to a deployable tracker that only uses past‑frame templates at inference. Experiments on seven benchmarks show consistent performance gains across both long‑ and short‑term tracking protocols.
By Xin Chen, Jiao Xu, Dong Wang, Huchuan Lu, Kede Ma
arXiv:2602. 14771v5 Announce Type: replace-cross Abstract: The human visual system tracks objects by integrating current observations with previously observed information, adapting to target and scene changes, and reasoning about occlusion at fine granularity.
By Shih-Fang Chen, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin
arXiv:2607. 01395v1 Announce Type: cross Abstract: At the heart of human visual perception lies the ability to maintain a continuous and coherent understanding of the external world.
By Shih-Fang Chen
arXiv:2609.22706v1 Announce Type: new
Abstract: Identity association in multi-object tracking (MOT) is vulnerable to partial occlusion, truncated detections, and fluctuating confidence scores. Existi...
By Hao Wang
arXiv:2610.01682v1 Announce Type: cross
Abstract: Mobile robots operating among pedestrians need trajectories that become available quickly, remain spatially credible through missed observations, pre...
By Dominik Wojcikiewicz, Diego Paez-Granados
arXiv:2609.24492v1 Announce Type: new
Abstract: Long-shot windsurfing video combines small targets, large camera pans, prolonged overlaps, and rapidly changing backgrounds. The desired output is not...
By Bertil Braun
GRACE is a camera‑efficient multi‑view pedestrian tracker that reduces the number of required cameras while maintaining high tracking accuracy. It combines volumetric‑guided fusion of homography‑based BEV features with 3D‑lifted features, uses ray conditioning to incorporate each camera’s viewing direction, and employs BEV Track Recovery to continue existing tracks with low‑confidence detections. On the WildTrack dataset, GRACE raises MOTA from 83.54 to 91.07 compared to the baseline TrackTacular.
By Taigo Sakai, Kazuhiro Hotta, Hiroki Kouno, Naoki Kato
Long-shot windsurfing video combines small targets, large camera pans, prolonged overlaps, and rapidly changing backgrounds. The desired output is not a generic MOT trace but a separate, stable rider-...
arXiv:2606. 23604v2 Announce Type: replace-cross Abstract: The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion estimation.
By Mohamed Nagy, Naoufel Werghi, Jorge Dias, Majid Khonji
arXiv:2609.18363v1 Announce Type: new
Abstract: Online multi camera 3D tracking must maintain scene global identities across synchronized views, yet query-based trackers carry these identities only i...
By Pragyan Shrestha, Haruto Nakayama, Atom Scott