arXiv AI

MGDFIS: Multi-scale Global-detail Feature Integration Strategy for Small Object Detection

arXiv:2506. 12697v3 Announce Type: replace-cross Abstract: Small-object detection in Unmanned Aerial Vehicle (UAV) imagery requires preserving weak local evidence while using broader context to separate tiny foreground targets from cluttered backgrounds.

Hugging Face Trending Papers
Jul 27

LCMamNet: A Lightweight Cross-scale Mamba Network for Infrared Small Target Detection

Infrared small target detection (IRSTD) is important for low-altitude perception, unmanned-system warning, and security monitoring. However, weak targets in infrared imagery usually occupy only a few pixels and are easily submerged by cloud clutter, ground edges, and bright noise, making it difficult for lightweight segmentation-based methods to preserve local target structures while suppressing background interference.

arXiv Computer Vision
Sep 17

PDA++: Field-Aligned Planning and Scene-Adaptive Insertion in Remote Sensing

PDA++ is a unified, environment‑aware object insertion framework for remote sensing imagery that improves few‑shot and long‑tail recognition. It operates in three stages: Planning, which selects scene‑compatible poses using an affordance field; Decoupling, which conditions the background on pose to preserve object identity while adapting to the scene; and Assimilation, which aligns multi‑scale texture distributions via optimal transport to enhance local coherence. The method achieves a whole‑image FID of 6.28 and boosts average few‑shot recognition mAP50 by 17.69 points on optical data, while also improving ship detection on SAR imagery and maintaining performance under cross‑dataset transfer and amorphous‑target insertion.

By Xianchi Dong, Yingyan Hou, Chao Ren, Wanxuan Lu, Zihan Wei, Hongfeng Yu, Yixiao Wang, Chubo Deng, Xian Sun
arXiv Computer Vision
Sep 17

Understanding Dynamic Scenes at Gigapixel Scale: Wide-Area Spatio-Temporal Perception from UAVs

The paper introduces the Wide-area Spatio-temporal Scene Understanding (WSTU) problem, which demands simultaneous wide-area coverage, per-target resolution, and temporal continuity—capabilities lacking in existing datasets. To address this, the authors present HARD, an ultra‑high‑resolution (12768×9564) UAV dataset annotated for object detection, multi‑object tracking, and scene‑level visual question answering. They also propose a latency‑aware metric, streaming‑HOTA (s‑HOTA), and show through baseline experiments that high resolution and processing latency significantly impact detection, tracking, and VQA performance, revealing gaps in current methods for WSTU.

By Yuhang Zhu, Meiyi Zhu, Yunkai Dang, Zhangnan Li, Yuxuan Wang, Wenbin Li, Hongbing Pan