arXiv AI By Wonbin Son, Gyumum Choi, Junil Seo, Hyungjoon Kim

What Frame-Level Labels Can and Cannot Do for Small-UAV Point Detection in Thermal Video

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computer Vision
Sep 3

RGB-to-IR image translation for infrared vehicle detection in unseen UAV domains

The paper explores using generative models to translate RGB UAV images into synthetic infrared (IR) images for training vehicle detectors in domains where real IR data is scarce. Various translators—supervised GANs, ControlNet-based diffusion models, and LoRA-ed foundation models—were trained on paired RGB-IR datasets and applied to unseen target datasets to generate synthetic IR data. The synthetic IR images, especially those produced by Stable Diffusion 3.5 with ControlNet, significantly improved detection performance on unseen IR test sets, outperforming RGB and grayscale baselines and narrowing the gap to real IR data.

By Thijs A. Eker, Ella P. Fokkinga, Jan Erik van Woerden, Elfi I. S. Hofmeijer, Sebastiaan P. Snel, Klamer Schutte, Friso G. Heslinga
arXiv AI
Aug 19

Training with synthetic data for drone detection in thermal imagery

The paper explores a synthetic-first training approach for detecting drones in medium- and long-wave infrared imagery, combining synthetic scene generation with fine-tuning on real data. It demonstrates that synthetic data can establish initial object representations, but real infrared data is crucial to close domain gaps and improve reliability. The study finds that aligning datasets has a greater impact on performance than increasing model size, and that semantic alignment in feature space is the strongest predictor of success, with radiometric factors like entropy and dynamic range also contributing.

By Tanel Liiv, Sander Soodla, Nzamba Bignoumba, Alma M. Liezenga, Toomas Pruuden
arXiv Computer Vision
Sep 17

Understanding Dynamic Scenes at Gigapixel Scale: Wide-Area Spatio-Temporal Perception from UAVs

The paper introduces the Wide-area Spatio-temporal Scene Understanding (WSTU) problem, which demands simultaneous wide-area coverage, per-target resolution, and temporal continuity—capabilities lacking in existing datasets. To address this, the authors present HARD, an ultra‑high‑resolution (12768×9564) UAV dataset annotated for object detection, multi‑object tracking, and scene‑level visual question answering. They also propose a latency‑aware metric, streaming‑HOTA (s‑HOTA), and show through baseline experiments that high resolution and processing latency significantly impact detection, tracking, and VQA performance, revealing gaps in current methods for WSTU.

By Yuhang Zhu, Meiyi Zhu, Yunkai Dang, Zhangnan Li, Yuxuan Wang, Wenbin Li, Hongbing Pan
arXiv Machine Learning
Aug 27

Model-Agnostic Open-Set Air-to-Air Visual Object Detection for Reliable UAV Perception

The paper introduces a model‑agnostic open‑set detection framework for air‑to‑air visual object detection on UAVs, addressing the limitations of closed‑set detectors under domain shifts and flight data corruption. It estimates semantic uncertainty through entropy modeling in the embedding space and employs spectral normalization and temperature scaling to improve open‑set discrimination. Experiments on the AOT aerial benchmark and real‑world flight tests show up to a 10% relative AUROC improvement over standard YOLO detectors, with background rejection further enhancing robustness without sacrificing accuracy.

By Spyridon Loukovitis, Anastasios Arsenos, Vasileios Karampinis, Athanasios Voulodimos