arXiv AI

VLA-ReID: Video-Level Association for Re-Identification in Multi-Object Tracking with Highly Similar Objects

arXiv:2607. 17157v1 Announce Type: cross Abstract: Multi-object tracking (MOT) aims to localize multiple objects in videos while preserving their identities over time.

arXiv Computer Vision
Sep 22

Beyond the Survey: A Systematic Empirical Study of Detection and Association in Visual MOT

This paper conducts a systematic empirical study of multi‑object tracking (MOT) algorithms, focusing on how detection and association components affect overall performance. By evaluating state‑of‑the‑art methods on benchmarks such as MOT16/17/20, SportsMOT, DanceTrack, and CrowdTrack, the authors find that detection quality has a far greater impact than association strategies, and that transformer‑based end‑to‑end models are more robust to detection variations but computationally expensive. The study provides a unified pipeline diagram and practical guidance for researchers and practitioners in selecting and designing MOT systems.

By Linh Van Ma, Juhua Hu, Wei Cheng, Unse Fatima, Moongu Jeon
Hugging Face Trending Papers
5d ago

VastMAT: A Large-Scale Multi-Category Benchmark for Multi-Animal Tracking

VastMAT is a large‑scale multi‑animal tracking benchmark featuring 2,947 videos, 337 animal categories, and over 3.6 million bounding boxes with 22,883 identity trajectories. It emphasizes high‑quality, expert‑reviewed annotations and introduces Seen‑category and Unseen‑category evaluation protocols, revealing significant challenges in tracking unseen animals. The authors also propose a lightweight Center‑Distance‑Augmented Association module that boosts HOTA scores for existing MOT methods without extra training.

arXiv Computer Vision
Sep 7

Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods

The paper introduces MovingDroneCrowd++, a large-scale video dataset for dense crowd counting and tracking from moving drones, featuring varied flight altitudes, camera angles, and lighting. It presents two new methods: GD3A for Video Individual Counting and GIA-Track for Multi-Object Tracking, both leveraging group-wise density assignment and identity association to handle aerial challenges. Experiments demonstrate significant improvements, reducing counting error by 47.4% and boosting tracking accuracy by 64.6%.

By Yaowu Fan, Jia Wan, Tao Han, Andy J. Ma, Wanli Ouyang, Antoni B. Chan
arXiv AI
3d ago

Template-Search Domain Adaptation via Multi-Stage Feature Alignment for Cross-Modal Object Tracking

The paper introduces TSDA-Track, a Template-Search Domain Adaptation framework designed to reduce modality gaps in cross‑modal visual object tracking. Two variants are explored: Pre‑AFA TSDA‑Track uses adversarial alignment before transformer interaction, while Enc‑CFA TSDA‑Track applies contrastive alignment after interaction to strengthen cross‑modal correspondence. Experiments on datasets such as LasHeR, RGBT234, GTOT, and Anti‑UAV‑024 show that both variants outperform state‑of‑the‑art trackers, with Pre‑AFA achieving an SR/PR of 43.2/56.0 on RGBT234 under the modality‑switch protocol.

By Fereshteh Aghaee Meibodi, Amir Mehdi Soufi Enayati, Shadi Alijani, Homayoun Najjaran