arXiv AI

Multi-Sensor Fusion for UAV Classification Based on Feature Maps of Image and Radar Data

arXiv:2410. 16089v2 Announce Type: replace Abstract: The unique cost, flexibility, speed, and efficiency of modern UAVs make them an attractive choice in many applications in contemporary society.

arXiv Computer Vision
Sep 2

Multimodal RGB-Infrared Combination for UAV-Based Wildfire Segmentation: A Comparative Study on FLAME3

The paper examines RGB‑infrared fusion for binary wildfire segmentation using UAV imagery on the FLAME3 dataset. It compares RGB and infrared baselines with three fusion strategies across U‑Net, DeepLabV3+, and SegFormer architectures. Results show thermal data dominates segmentation performance, and feature‑level fusion with transformer‑based models yields the best results.

By Matheus F. Kovaleski, Lu\'is Garrote, Cristiano Premebida, J\'er\^ome Mendes, Jo\~ao Ruivo Paulo
Hugging Face Trending Papers
Jun 4

Comparison of Deep Learning Frameworks For Rice Disease Mapping From UAV Multispectral Imaging

In this study, UAV multispectral imagery is used to segment the severity of bacterial leaf blight (BLB) in rice using convolutional neural networks (CNNs) and transformer-based models. The evaluated architectures include U-Net with a ResNet- 101 encoder, U-Net++ with EfficientNet-B3 and EfficientNetB7, DeepLabV3+, and SegFormer, all trained under a common pipeline with three input configurations (multispectral only, multispectral+NDVI, and multispectral+NDRE).

arXiv Computer Vision
Sep 3

Improved Automatic Target Recognition in Synthetic Aperture Sonar Imagery Using Large Deep Neural Networks

The paper investigates Automatic Target Recognition (ATR) in Synthetic Aperture Sonar (SAS) imagery, comparing modern convolutional neural networks (CNNs) and transformer-based deep neural networks (DNNs). It examines how factors such as network size, architecture, pretraining methods, data augmentation, and regularization influence performance, aiming to identify the highest-performing model and provide a training roadmap for state‑of‑the‑art SAS‑ATR systems.

By C. J. Moore, Alex Hurt, Jordan Malof
arXiv Computer Vision
Sep 3

Evidential Deep Learning for Multi-Modal Anti-UAV Detection

The paper investigates the use of evidential deep learning (EDL) for multi‑modal anti‑UAV detection, comparing it with sigmoid baselines, Dempster‑Shafer evidence fusion, and uncertainty‑driven temporal sensor gating across three benchmarks (thermal tracking, RGB‑audio‑RF classification, and RGB‑IR tracking). EDL improves accuracy by up to 5.9 percentage points and better ranks classification errors, while the other components (DS fusion, Dirichlet vacuity, temporal gating) do not provide the expected benefits. The study concludes that the primary advantage of EDL stems from its training objective rather than its uncertainty estimates.

By Dmitry Golovchits, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag
arXiv Computer Vision
Aug 28

UFO-DETR: Frequency-Guided End-to-End Detector for UAV Tiny Objects

UFO-DETR is an end‑to‑end object detection framework designed for UAV imagery, addressing challenges such as scale variation, dense distribution, and the prevalence of tiny targets. It employs an LSKNet backbone to optimize receptive fields and reduce parameters, integrates DAttention and AIFI modules for flexible multi‑scale spatial modeling, and introduces a DynFreq‑C3 module that enhances small target detection via cross‑space frequency feature enhancement. Experiments demonstrate that UFO‑DETR outperforms RT‑DETR‑L in detection accuracy while improving computational efficiency, making it suitable for UAV edge computing.

By Yuankai Chen, Kai Lin, Qihong Wu, Xinxuan Yang, Jiashuo Lai, Ruoen Chen, Haonan Shi, Minfan He, Meihua Wang
arXiv Machine Learning
Sep 17

The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification

The paper investigates how to balance model size and fine‑tuning strategy for UAV audio classification. Using a dataset of 3,100 clips across 31 drone classes, it compares transformer and convolutional backbones under full fine‑tuning, classifier‑only fine‑tuning, and four parameter‑efficient fine‑tuning methods. Results show that selective batch‑norm tuning of EfficientNet‑B7 yields the best accuracy (97.65%) while updating less than 0.5% of parameters, and that lightweight CNNs generally outperform transformers in both accuracy and efficiency.

By Andrew P. Berg, Qian Zhang, Mia Y. Wang