arXiv Computer Vision

EECTracker: Swarm Motion Prior-Guided Feature Compensation for Airborne Optical UAV Swarm Tracking

arXiv Computer Vision
Sep 7

Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods

The paper introduces MovingDroneCrowd++, a large-scale video dataset for dense crowd counting and tracking from moving drones, featuring varied flight altitudes, camera angles, and lighting. It presents two new methods: GD3A for Video Individual Counting and GIA-Track for Multi-Object Tracking, both leveraging group-wise density assignment and identity association to handle aerial challenges. Experiments demonstrate significant improvements, reducing counting error by 47.4% and boosting tracking accuracy by 64.6%.

By Yaowu Fan, Jia Wan, Tao Han, Andy J. Ma, Wanli Ouyang, Antoni B. Chan
arXiv Computer Vision
Sep 17

Understanding Dynamic Scenes at Gigapixel Scale: Wide-Area Spatio-Temporal Perception from UAVs

The paper introduces the Wide-area Spatio-temporal Scene Understanding (WSTU) problem, which demands simultaneous wide-area coverage, per-target resolution, and temporal continuity—capabilities lacking in existing datasets. To address this, the authors present HARD, an ultra‑high‑resolution (12768×9564) UAV dataset annotated for object detection, multi‑object tracking, and scene‑level visual question answering. They also propose a latency‑aware metric, streaming‑HOTA (s‑HOTA), and show through baseline experiments that high resolution and processing latency significantly impact detection, tracking, and VQA performance, revealing gaps in current methods for WSTU.

By Yuhang Zhu, Meiyi Zhu, Yunkai Dang, Zhangnan Li, Yuxuan Wang, Wenbin Li, Hongbing Pan
arXiv AI
2d ago

Template-Search Domain Adaptation via Multi-Stage Feature Alignment for Cross-Modal Object Tracking

The paper introduces TSDA-Track, a Template-Search Domain Adaptation framework designed to reduce modality gaps in cross‑modal visual object tracking. Two variants are explored: Pre‑AFA TSDA‑Track uses adversarial alignment before transformer interaction, while Enc‑CFA TSDA‑Track applies contrastive alignment after interaction to strengthen cross‑modal correspondence. Experiments on datasets such as LasHeR, RGBT234, GTOT, and Anti‑UAV‑024 show that both variants outperform state‑of‑the‑art trackers, with Pre‑AFA achieving an SR/PR of 43.2/56.0 on RGBT234 under the modality‑switch protocol.

By Fereshteh Aghaee Meibodi, Amir Mehdi Soufi Enayati, Shadi Alijani, Homayoun Najjaran
arXiv AI
Aug 11

RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation

arXiv:2608. 09467v1 Announce Type: cross Abstract: Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments.

By Boxiong Wang, Hui Kang, Geng Sun, Jiahui Li, Chao Yu, Daxin Tian
arXiv Computer Vision
Sep 3

Evidential Deep Learning for Multi-Modal Anti-UAV Detection

The paper investigates the use of evidential deep learning (EDL) for multi‑modal anti‑UAV detection, comparing it with sigmoid baselines, Dempster‑Shafer evidence fusion, and uncertainty‑driven temporal sensor gating across three benchmarks (thermal tracking, RGB‑audio‑RF classification, and RGB‑IR tracking). EDL improves accuracy by up to 5.9 percentage points and better ranks classification errors, while the other components (DS fusion, Dirichlet vacuity, temporal gating) do not provide the expected benefits. The study concludes that the primary advantage of EDL stems from its training objective rather than its uncertainty estimates.

By Dmitry Golovchits, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag
arXiv Computer Vision
Aug 27

STATrack: A Target-Aware Fully Spiking Neural Network for Efficient UAV Tracking

STATrack is a fully spiking neural network designed for UAV visual tracking using only RGB inputs, eliminating the need for costly event cameras. It introduces Adaptive Mutual Information Maximization (AMIM) to preserve fine-grained target information in deep spiking representations and a sample-difficulty-aware dynamic weighting strategy to adjust the mutual‑information constraint during training. Experiments on four UAV tracking benchmarks show that STATrack achieves state‑of‑the‑art performance while maintaining low theoretical energy consumption.

By Pengzhi Zhong, Jiwei Mo, Dan Zeng, Feixiang He, Shuiwang Li