The ByteTrack algorithm is a widely used and computationally efficient multi-object tracking architecture. Its core innovation lies in the combination of lenient bounding box associations with trackle...
This paper conducts a systematic empirical study of multi‑object tracking (MOT) algorithms, focusing on how detection and association components affect overall performance. By evaluating state‑of‑the‑art methods on benchmarks such as MOT16/17/20, SportsMOT, DanceTrack, and CrowdTrack, the authors find that detection quality has a far greater impact than association strategies, and that transformer‑based end‑to‑end models are more robust to detection variations but computationally expensive. The study provides a unified pipeline diagram and practical guidance for researchers and practitioners in selecting and designing MOT systems.
By Linh Van Ma, Juhua Hu, Wei Cheng, Unse Fatima, Moongu Jeon
arXiv:2609.12261v1 Announce Type: new
Abstract: Multi-object tracking (MOT) is dominated by the tracking-by-detection paradigm, whose methods typically rely on a small set of hyperparameters that are...
By Momir Ad\v{z}emovi\'c
arXiv:2606. 23604v2 Announce Type: replace-cross Abstract: The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion estimation.
By Mohamed Nagy, Naoufel Werghi, Jorge Dias, Majid Khonji
arXiv:2607. 17157v1 Announce Type: cross Abstract: Multi-object tracking (MOT) aims to localize multiple objects in videos while preserving their identities over time.
By Yanrong Qin, Xiaoyan Cao, Yao Yao
Tetris is a video object tracking system that uses tile-level sampling to efficiently extract high‑fidelity tracks. It partitions videos into tile‑based polyominoes, classifies relevant tiles, prunes redundant ones with an ILP under a user‑defined accuracy constraint, and packs the remaining polyominoes to minimize detector calls. On seven stationary‑video datasets, Tetris maintains less than a 5% loss in tracking accuracy while achieving up to 17.4× higher throughput than prior systems and up to 68.8× higher than a full‑frame reference pipeline.
By Chanwut Kittivorawong, Alena Chao, Charlie Si, Alvin Cheung