arXiv:2512. 18046v2 Announce Type: replace Abstract: Unmanned Aerial Vehicles, commonly known as, drones pose increasing risks in civilian and defense settings, demanding accurate and real-time drone detection systems.
By Ami Pandat, Punna Rajasekhar, Gopika Vinod, Rohit Shukla
General-purpose object detectors lose accuracy on UAV footage, where targets span only a handful of pixels and onboard compute is limited. Prior work composes independently-validated architectural tec...
arXiv:2609.14560v1 Announce Type: new
Abstract: General-purpose object detectors lose accuracy on UAV footage, where targets span only a handful of pixels and onboard compute is limited. Prior work c...
By Quratulain Nayeem, Fahmina Taranum, Mohammed Mudassir Uddin
The paper introduces Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers high‑resolution spatial representations from a YOLO11m‑P2 teacher to a lightweight YOLO11n student without changing the student’s inference architecture. CSCWD aligns teacher P2 features with student P3 while also applying same‑scale distillation at deeper pyramid levels, yielding a 2.92‑point mAP@0.5 improvement over the baseline and a 2.09‑point gain over same‑scale distillation alone. In zero‑shot tests on DUT‑Anti‑UAV and on a Raspberry Pi 5, the 2.58‑million‑parameter student reaches 50.32% mAP@0.5 at 82.32 ms latency (12.15 fps) with negligible runtime or memory increase.
By Amir Zamani, Zeinab Ghasemi-Naraghi
HGSQ is a real‑time aerial small‑object detector that uses a heatmap‑guided sparse query strategy to focus computation on foreground regions. It introduces a lightweight Heatmap Budget Predictor to generate a foreground budget map, and then employs Heatmap‑Guided Sparse Query Selection, Heatmap‑Gated Lite Snake Convolution, and Adaptive Query‑Decoder Budgeting to efficiently process only small‑object areas. On NWPU VHR‑10 and VisDrone2019, HGSQ achieves 95.10 mAP50 and 54.8 mAP50 respectively while running at 96 FPS with only 48.6 GFLOPs on an RTX 4070.
By Yangchen Zeng
arXiv:2409. 16808v3 Announce Type: replace-cross Abstract: Modern applications such as autonomous vehicles, intelligent surveillance, and smart city systems increasingly require object detection on resource-constrained edge devices.
By Daghash K. Alqahtani, Muhammad Aamir Cheema, Maria A. Rodriguez, Adel N. Toosi
arXiv:2609.13647v1 Announce Type: new
Abstract: The rapid development of unmanned aerial vehicle (UAV) technology has made aerial-image object detection increasingly important for natural-resource mo...
By Hao Wang
arXiv:2609.23061v1 Announce Type: new
Abstract: Small-object detection in UAV imagery is challenged by weak visual evidence, ambiguous boundaries, dense object distributions, and complex backgrounds....
By Linduo Wei, Junjie Fan, Yijun Mai, Yong Qi
DenseScout is a 1.01M‑parameter dense‑response selector designed for edge platforms that directly optimizes ranked patch‑center prioritization, eliminating the need for detector‑style box regression. It aligns output representation, supervision, and decoding, and is jointly designed with transport‑aware execution and QoS‑oriented evaluation. Experiments on VisDrone and DOTA show that DenseScout achieves stronger low‑budget recall than detector‑derived selectors, and cross‑platform profiling on Jetson Orin NX and RK3588 demonstrates that deployable utility depends on selector quality, memory movement, and heterogeneous runtime realization.
By Zhouzhi Xiong, Zimo Zeng, Yi Chen, Shuqi Xu, Yunfeng Yan, Donglian Qi
arXiv:2609.10156v1 Announce Type: new
Abstract: Small object detection in unmanned aerial vehicle (UAV) and remote sensing imagery requires preserving high-resolution detail while modeling long-range...
By Junjie Fan, Yijun Mai, Linduo Wei, Jiayu Rao, Junmin Bao, Qiushi Jin, Guijia Li, Yong Qi
The paper evaluates the claim that vision‑language models (VLMs) outperform task‑specific vision backbones for UAV power‑line defect assessment using the ElecVQA‑Bench benchmark. Across various evaluation settings—partitioning, item sets, label spaces, replication, resolution, and side information—the performance gap between VLMs and traditional backbones is minimal or even reversed when controlling for resolution and token budget. The study concludes that VLM superiority is not universally supported and emphasizes the importance of rigorous benchmark audits.
By Linghao Zhang, Siyu Xiang, Junwei Kuang, Peiyu Yi
arXiv:2608.28706v1 Announce Type: new
Abstract: ViT detectors fix a uniform token grid before any learned stage. A native-resolution aerial detector must then choose between resolving few-pixel objec...
By Khayrul Islam