arXiv Computer Vision

ByteTraX: Enhancing the ByteTrack Architecture with Optimised Thresholding

ByteTraX is a lightweight enhancement to the ByteTrack multi‑object tracking architecture that introduces a single unified matching threshold and stricter track initiation criteria to reduce erroneous track reclassification and identity switches. The method yields consistent performance gains across several benchmarks—GMOT‑40, LC‑MOT, SportsMOT, TeamTrack, DAMUNT, and DeepSea‑MOT—while boosting processing speed by over 10%. Quantitatively, ByteTraX achieves more than a 40% drop in identity switches, with mean improvements of 3.6 in HOTA, 5.6 in IDF1, and 6.3 FPS.

arXiv Computer Vision
Sep 22

Beyond the Survey: A Systematic Empirical Study of Detection and Association in Visual MOT

This paper conducts a systematic empirical study of multi‑object tracking (MOT) algorithms, focusing on how detection and association components affect overall performance. By evaluating state‑of‑the‑art methods on benchmarks such as MOT16/17/20, SportsMOT, DanceTrack, and CrowdTrack, the authors find that detection quality has a far greater impact than association strategies, and that transformer‑based end‑to‑end models are more robust to detection variations but computationally expensive. The study provides a unified pipeline diagram and practical guidance for researchers and practitioners in selecting and designing MOT systems.

By Linh Van Ma, Juhua Hu, Wei Cheng, Unse Fatima, Moongu Jeon
arXiv Computer Vision
Aug 24

Tetris: Tile-level Sampling for Efficient and High-Fidelity Video Object Tracking

Tetris is a video object tracking system that uses tile-level sampling to efficiently extract high‑fidelity tracks. It partitions videos into tile‑based polyominoes, classifies relevant tiles, prunes redundant ones with an ILP under a user‑defined accuracy constraint, and packs the remaining polyominoes to minimize detector calls. On seven stationary‑video datasets, Tetris maintains less than a 5% loss in tracking accuracy while achieving up to 17.4× higher throughput than prior systems and up to 68.8× higher than a full‑frame reference pipeline.

By Chanwut Kittivorawong, Alena Chao, Charlie Si, Alvin Cheung
Hugging Face Trending Papers
4d ago

VastMAT: A Large-Scale Multi-Category Benchmark for Multi-Animal Tracking

VastMAT is a large‑scale multi‑animal tracking benchmark featuring 2,947 videos, 337 animal categories, and over 3.6 million bounding boxes with 22,883 identity trajectories. It emphasizes high‑quality, expert‑reviewed annotations and introduces Seen‑category and Unseen‑category evaluation protocols, revealing significant challenges in tracking unseen animals. The authors also propose a lightweight Center‑Distance‑Augmented Association module that boosts HOTA scores for existing MOT methods without extra training.