arXiv Computer Vision By Thomas A. O'Shea-Wheller

ByteTraX: Enhancing the ByteTrack Architecture with Optimised Thresholding

Read the original on arXiv Computer Vision →

ByteTraX is a lightweight enhancement to the ByteTrack multi‑object tracking architecture that introduces a single unified matching threshold and stricter track initiation criteria to reduce erroneous track reclassification and identity switches. The method yields consistent performance gains across several benchmarks—GMOT‑40, LC‑MOT, SportsMOT, TeamTrack, DAMUNT, and DeepSea‑MOT—while boosting processing speed by over 10%. Quantitatively, ByteTraX achieves more than a 40% drop in identity switches, with mean improvements of 3.6 in HOTA, 5.6 in IDF1, and 6.3 FPS.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 22

Beyond the Survey: A Systematic Empirical Study of Detection and Association in Visual MOT

This paper conducts a systematic empirical study of multi‑object tracking (MOT) algorithms, focusing on how detection and association components affect overall performance. By evaluating state‑of‑the‑art methods on benchmarks such as MOT16/17/20, SportsMOT, DanceTrack, and CrowdTrack, the authors find that detection quality has a far greater impact than association strategies, and that transformer‑based end‑to‑end models are more robust to detection variations but computationally expensive. The study provides a unified pipeline diagram and practical guidance for researchers and practitioners in selecting and designing MOT systems.

By Linh Van Ma, Juhua Hu, Wei Cheng, Unse Fatima, Moongu Jeon
arXiv Computer Vision
Aug 24

Tetris: Tile-level Sampling for Efficient and High-Fidelity Video Object Tracking

Tetris is a video object tracking system that uses tile-level sampling to efficiently extract high‑fidelity tracks. It partitions videos into tile‑based polyominoes, classifies relevant tiles, prunes redundant ones with an ILP under a user‑defined accuracy constraint, and packs the remaining polyominoes to minimize detector calls. On seven stationary‑video datasets, Tetris maintains less than a 5% loss in tracking accuracy while achieving up to 17.4× higher throughput than prior systems and up to 68.8× higher than a full‑frame reference pipeline.

By Chanwut Kittivorawong, Alena Chao, Charlie Si, Alvin Cheung