arXiv:2607. 02718v1 Announce Type: cross Abstract: Recent advances in large-scale image generative models enable photorealistic scene synthesis with controllable attributes.
By Stanislav Panev, Minhyek Jeon, Vaishnavi Khindkar, Ahish Deshpande, Celso M de Melo, Shuowen Hu, Shayok Chakraborty, Fernando De la Torre
The paper explores a synthetic-first training approach for detecting drones in medium- and long-wave infrared imagery, combining synthetic scene generation with fine-tuning on real data. It demonstrates that synthetic data can establish initial object representations, but real infrared data is crucial to close domain gaps and improve reliability. The study finds that aligning datasets has a greater impact on performance than increasing model size, and that semantic alignment in feature space is the strongest predictor of success, with radiometric factors like entropy and dynamic range also contributing.
By Tanel Liiv, Sander Soodla, Nzamba Bignoumba, Alma M. Liezenga, Toomas Pruuden
arXiv:2608.21926v1 Announce Type: new
Abstract: Unmanned aerial vehicle (UAV) navigation in modern low-altitude environments requires more accurate pose alignment in the final approach stage for targ...
By Jinyi Zhou, Shuo Feng, Yufei Wu, Piji Li
arXiv:2511.15343v2 Announce Type: replace-cross
Abstract: Autonomous navigation in complex scenes requires reliable perception across scenarios that the model did not encounter during its training. A...
By Spyridon Loukovitis, Vasileios Karampinis, Athanasios Voulodimos
arXiv:2606. 00747v1 Announce Type: cross Abstract: For low-altitude Unmanned Aerial Vehicle (UAV) autonomy, 3D spatial understanding is not merely a perception objective, but the safety interface between human instructions and physical flight.
By Jie Gao, Jie Ma, Kaihui Lin, Kai Ye, Miaohui Zhang, Pingyang Dai, Liujuan Cao
arXiv:2608.27214v1 Announce Type: new
Abstract: Open World Object Detection (OWOD) built on multimodal foundation models often suffers from semantic ambiguity caused by unidirectional text-to-vision...
By Hao Xu, Zhaoning Shi, Hehe Jin, Bo Ma
Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large-scale outdoor environments with substantial future-state uncertainty.
arXiv:2608. 09564v1 Announce Type: cross Abstract: UAV vision-language navigation (UAV-VLN) focuses on enabling an aerial agent to follow natural-language instructions in open 3D environments from egocentric visual observations.
By Zeyuan Ma, Jiaxin Chen, Di Huang
The paper introduces a saliency-depth conditioning approach for zero‑shot segmentation of communication‑tower components in cluttered UAV imagery. By combining appearance‑based saliency with monocular relative depth, the method creates a coarse tower prior that suppresses irrelevant background, and integrates this module with Grounded‑SAM and SAM 3 to produce SD‑Grounded‑SAM and SD‑SAM 3. Experiments on the TOW‑300 dataset show that SD‑SAM 3 achieves the best instance‑segmentation performance while SD‑Grounded‑SAM reduces false positives, with ablations confirming the complementary benefits of saliency, depth, and box refinement.
By Ali Lesani, Chul Min Yeum, Su-Min Kang
arXiv:2608. 08308v1 Announce Type: cross Abstract: Modern vision systems must operate in "open-world" settings, where models must recognize known categories and detect unseen or anomalous content.
By Anastasios Romanos Varvarigos, Nikos Giakoumoglou, Tania Stathaki
arXiv:2606. 15468v1 Announce Type: cross Abstract: Vision models can achieve strong performance on classification tasks, but the internal representations supporting their predictions are often difficult to interpret.
By Deepshik Sharma
arXiv:2606. 24353v1 Announce Type: cross Abstract: Bird's-eye view (BEV) perception fuses multi-camera images into a unified top-down representation for autonomous driving.
By Hojun Choi, Seulbin Hwang, Dae Jung Kim, Kisung Kim, Hyunjung Shim, Jinhan Lee