arXiv Computer Vision By Markus Gross, Andreas Greiner, Taehyoung Kim, Sivasubiramaniam Subbiah, Toma\v{z} Coti\v{c}, Sai Bharadwaj Matha, Conrad Christoph, Oussema Dhaouadi, Simon Zieher, Surya Vijaya Kumar, Gordon Elger, Henri Mee{\ss}, Olaf Wysocki, Paul Spannaus, Daniel Cremers

MultiFly: A Real-World Multimodal Aerial Dataset with Annotation-Efficient Label Transfer and Cross-Modal Semantic Consistency

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

arXiv AI
Sep 30

AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

AerialDojo-200K is a large-scale benchmark suite for open-world aerial object-goal search, featuring 42 simulation scenes across four families and 21 types, including urban, natural, infrastructure, and disaster environments. The dataset contains 205,732 task instances—over 100K semantic-goal and over 100K image-goal tasks—each with a collision-free reference trajectory and multi-view video recordings. A unified evaluation framework splits scenes into 21 in-distribution and 21 out-of-distribution sets, and preliminary tests on multimodal large language models show significant room for improvement in general-purpose aerial agents.

By Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang, Weiran Wang, Shaokai Zhu, Ziqi Jia, Hao Wang, Yu-Wei Zhan, Zongyuan Wu, Jinghao Cui, Wenwu Zhu
arXiv Computer Vision
Sep 22

M3GA-Wild: A Large-Scale Dataset and Benchmark for Multi-Modal Multi-session Ground-to-Aerial Place Recognition in Forests

M3GA-Wild is a new benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests, combining synchronized RGB imagery and LiDAR from ground traversals with high‑resolution aerial imagery and multi‑altitude LiDAR over 370 hectares. The dataset includes accurate geo‑referenced 6‑DoF poses and spans 36 km of forest traversals, enabling systematic evaluation of visual, LiDAR, cross‑modal, and multi‑modal methods. Baseline experiments show LiDAR outperforms vision‑only approaches under severe viewpoint changes, while current multi‑modal fusion offers limited gains due to poor cross‑modal alignment, highlighting challenges in cross‑platform localisation and domain gaps.

By Ethan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes, Milad Ramezani
arXiv Computer Vision
Sep 1

CedarCypress3D: an annotated UAV-LiDAR dataset of individual trees in planted cedar and cypress forests

arXiv:2608.30149v1 Announce Type: new Abstract: Individual tree measurements derived from Light Detection and Ranging (LiDAR) mounted on Unmanned Aerial Vehicles (UAV) provide valuable information fo...

By Katsuto Shimizu (Shikoku Research Center, Forestry and Forest Products Research Institute), Fumiaki Kitahara (Department of Forest Management, Forestry and Forest Products Research Institute), Tomohiro Nishizono (Department of Forest Management, Forestry and Forest Products Research Institute), Hideki Saito (Forestry and Forest Products Research Institute), Masayoshi Takahashi (Department of Forest Management, Forestry and Forest Products Research Institute), Shingo Obata (Hokkaido Research Center, Forestry and Forest Products Research Institute), Shunsuke Tei (Hokkaido Research Center, Forestry and Forest Products Research Institute), Naoyuki Furuya (Department of Forest Management, Forestry and Forest Products Research Institute), Tomoya Goto (Green Kogyo), Eiji Kodani (Department of Forest Management, Forestry and Forest Products Research Institute), Yusuke Yamada (Graduate School of Bioagricultural Sciences, Nagoya University)