arXiv:2609.26549v1 Announce Type: new
Abstract: Individual tree crown segmentation from aerial imagery underpins tree-level carbon accounting, biodiversity, and restoration monitoring at landscape sc...
By Thomas Pitts, Kunqi Li, Bin Liang
Infrared small target detection (IRSTD) is important for low-altitude perception, unmanned-system warning, and security monitoring. However, weak targets in infrared imagery usually occupy only a few pixels and are easily submerged by cloud clutter, ground edges, and bright noise, making it difficult for lightweight segmentation-based methods to preserve local target structures while suppressing background interference.
arXiv:2608.28706v1 Announce Type: new
Abstract: ViT detectors fix a uniform token grid before any learned stage. A native-resolution aerial detector must then choose between resolving few-pixel objec...
By Khayrul Islam
arXiv:2608. 15790v1 Announce Type: new Abstract: Crevasse mapping from uncrewed aerial vehicle (UAV) imagery matters for glaciological research and for field safety in glaciated terrain.
By Steven Wallace, William D Harcourt, Richard Hann, Aiden Durrant, Somayajulu Sripada, Georgios Leontidis
arXiv:2506. 12697v3 Announce Type: replace-cross Abstract: Small-object detection in Unmanned Aerial Vehicle (UAV) imagery requires preserving weak local evidence while using broader context to separate tiny foreground targets from cluttered backgrounds.
By Yuxiang Wang, Xuecheng Bai, Chuanzhi Xu, Ying Zhou, Weidong Cai
The paper evaluates out‑of‑the‑box object detection models for automatic target detection and recognition (ATD/R) in military settings. Six YOLO variants and two DETR variants were benchmarked on a new military dataset featuring vehicles, occlusions, and small targets, with performance measured in mAP@0.5 and mAP@0.5:0.95 across air‑to‑ground and ground‑to‑ground perspectives. Findings show larger models and DETR-based approaches perform best, fine‑tuning on the VisDrone dataset improves air‑to‑ground and small‑object performance, yet all models still struggle with small targets in air‑to‑ground scenarios.
By Alma M. Liezenga, Lotte Nijskens, Henrik R. Baumann, Stefan Becker, Simon Bensberg, Niccol\`o Camarlinghi, H{\aa}vard R. Eiring, Alexander W. Johnsgaard, Tanel Liiv, Giuseppe Martino, Matteo Marturini, Matthias Rapp, Jan Erik van Woerden, Alexander Wolpert, Hugo J. Kuijf
SAM‑V is a geometry‑aware extension of the Segment Anything Model (SAM) that integrates 3D priors from a feed‑forward geometry model (VGGT) into 2D segmentation. It uses a prompt‑fusion mechanism to combine sparse SAM prompts with view‑specific camera tokens and local VGGT features, enabling a mask decoder that attends to both dense 2D and 3D cues. The resulting end‑to‑end system produces consistent multi‑view instance segmentation in a single forward pass, achieving significant gains on the IGGT 3D tracking benchmark without offline mask matching or explicit 3D reconstruction.
By Jiangshan Gong, Yuqun Wu, Qiqian Fu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem
The paper introduces a saliency-depth conditioning approach for zero‑shot segmentation of communication‑tower components in cluttered UAV imagery. By combining appearance‑based saliency with monocular relative depth, the method creates a coarse tower prior that suppresses irrelevant background, and integrates this module with Grounded‑SAM and SAM 3 to produce SD‑Grounded‑SAM and SD‑SAM 3. Experiments on the TOW‑300 dataset show that SD‑SAM 3 achieves the best instance‑segmentation performance while SD‑Grounded‑SAM reduces false positives, with ablations confirming the complementary benefits of saliency, depth, and box refinement.
By Ali Lesani, Chul Min Yeum, Su-Min Kang
General-purpose object detectors lose accuracy on UAV footage, where targets span only a handful of pixels and onboard compute is limited. Prior work composes independently-validated architectural tec...
The paper presents ViTok, a multi‑teacher distillation approach that combines SigLIP2 and DINOv3‑L to jointly preserve global recognition and dense semantics. By introducing split adaptor heads, asymmetric losses, teacher reweighting, masked image modeling, and PHI‑S feature balancing, the authors achieve higher ImageNet‑1K kNN accuracy and restore ADE20K segmentation performance to match the teacher. The study also reports negative findings, such as limited benefits from scaling to ImageNet22K and interference from additional teachers.
By Hailun Xu, Kanchan Sarkar
arXiv:2603.15812v3 Announce Type: replace
Abstract: Multi-View Multi-Object Tracking (MV-MOT) aims to localize and maintain consistent identities of objects observed by multiple sensors. This task is...
By Aditya Iyer, Jack Roberts, Nora Ayanian
arXiv:2609.14560v1 Announce Type: new
Abstract: General-purpose object detectors lose accuracy on UAV footage, where targets span only a handful of pixels and onboard compute is limited. Prior work c...
By Quratulain Nayeem, Fahmina Taranum, Mohammed Mudassir Uddin