SiST‑GNN introduces a simultaneous spatial‑temporal message‑passing framework for dynamic graph neural networks, fusing per‑node temporal embeddings with spatial aggregation in a single operation. By maintaining a recurrent hidden state per node and treating it as a cross‑time edge, the model jointly reasons over topology and evolution. Experiments on link‑prediction and node‑classification benchmarks show significant improvements over prior methods, achieving up to 158% gains in live‑update link prediction and outperforming discrete‑time baselines by 7–23% in dynamic node classification.
By Shubhajit Roy, Anirban Dasgupta
This paper conducts a systematic empirical study of multi‑object tracking (MOT) algorithms, focusing on how detection and association components affect overall performance. By evaluating state‑of‑the‑art methods on benchmarks such as MOT16/17/20, SportsMOT, DanceTrack, and CrowdTrack, the authors find that detection quality has a far greater impact than association strategies, and that transformer‑based end‑to‑end models are more robust to detection variations but computationally expensive. The study provides a unified pipeline diagram and practical guidance for researchers and practitioners in selecting and designing MOT systems.
By Linh Van Ma, Juhua Hu, Wei Cheng, Unse Fatima, Moongu Jeon
arXiv:2603. 24016v2 Announce Type: replace-cross Abstract: Multi-Object Tracking (MOT) has traditionally focused on a few specific categories, restricting its applicability to real-world scenarios involving diverse objects.
By Zekun Qian, Wei Feng, Ruize Han, Junhui Hou
arXiv:2608. 16142v1 Announce Type: cross Abstract: UAV on-board vision systems are widely used for different activities, including monitoring in no-fly zones.
By Alam Noor, Luis Almeida, Kai Li, Jiyan Wu, Miguel Guti\'errez Gait\'an, Eduardo Tovar
The paper introduces the Wide-area Spatio-temporal Scene Understanding (WSTU) problem, which demands simultaneous wide-area coverage, per-target resolution, and temporal continuity—capabilities lacking in existing datasets. To address this, the authors present HARD, an ultra‑high‑resolution (12768×9564) UAV dataset annotated for object detection, multi‑object tracking, and scene‑level visual question answering. They also propose a latency‑aware metric, streaming‑HOTA (s‑HOTA), and show through baseline experiments that high resolution and processing latency significantly impact detection, tracking, and VQA performance, revealing gaps in current methods for WSTU.
By Yuhang Zhu, Meiyi Zhu, Yunkai Dang, Zhangnan Li, Yuxuan Wang, Wenbin Li, Hongbing Pan
arXiv:2609.22706v1 Announce Type: new
Abstract: Identity association in multi-object tracking (MOT) is vulnerable to partial occlusion, truncated detections, and fluctuating confidence scores. Existi...
By Hao Wang
arXiv:2407.04308v4 Announce Type: replace-cross
Abstract: We propose a graph-based tracking formulation for multi-object tracking (MOT) where target detections contain kinematic information and re-id...
By Griffin Golias, Masa Nakura-Fan, Vitaly Ablavsky
VastMAT is a large‑scale multi‑animal tracking benchmark featuring 2,947 videos, 337 animal categories, and over 3.6 million bounding boxes with 22,883 identity trajectories. It emphasizes high‑quality, expert‑reviewed annotations and introduces Seen‑category and Unseen‑category evaluation protocols, revealing significant challenges in tracking unseen animals. The authors also propose a lightweight Center‑Distance‑Augmented Association module that boosts HOTA scores for existing MOT methods without extra training.
UAV on-board vision systems are widely used for different activities, including monitoring in no-fly zones. In this case, the vision-equipped UAV streams a video to a ground server where an operator assists its activities.
SBMVTrack is a fully spiking neural network framework designed for energy-efficient UAV visual tracking. It introduces Energy-Weighted Spike Budgeting (EWSB) to constrain spike activity based on computational cost, and Masked Multi-View Target Modeling (MVTM) to enhance target representation by leveraging correlated temporal views. Experiments on multiple benchmarks show that SBMVTrack reduces spike firing rates and theoretical energy consumption while maintaining competitive tracking accuracy.
By Pengzhi Zhong, Jiwei Mo, Haolun Li, Ge Zheng, Jingqi Wang, Xinyi Bo, Shuiwang Li
arXiv:2410. 19553v2 Announce Type: replace-cross Abstract: This paper explores the impact of occlusions in video action detection.
By Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat
arXiv:2511. 18493v4 Announce Type: replace-cross Abstract: The significant variability in cell size and shape continues to pose a major obstacle in computer-assisted cancer detection on gigapixel Whole Slide Images (WSIs), due to cellular heterogeneity.
By Gia Huy Thai, Hoang-Nguyen Vu, Anh-Minh Phan, Quang-Thinh Ly, Thi-Ngoc-Truc Nguyen, Nhat Ho