LTM: Large-scale Terrain Model for Wildfire-prone Landscapes
arXiv:2607. 08711v1 Announce Type: cross Abstract: Accurate 3D terrain maps are essential for emergency response when assessing wildfire hazards.
Accurate 3D terrain maps are essential for emergency response when assessing wildfire hazards. However, wildfire-prone regions often span vast areas where conventional reconstruction methods underperform.
arXiv:2607. 08711v1 Announce Type: cross Abstract: Accurate 3D terrain maps are essential for emergency response when assessing wildfire hazards.
We present M3GA-Wild, the first benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests. M3GA-Wild unifies and extends existing forest localisation datasets, providing a...
arXiv:2608.22821v1 Announce Type: new Abstract: We present SiZeUp, a fast and scalable approach for constructing large-scale 3D urban proxy models directly from calibrated oblique aerial imagery. Our...
M3GA-Wild is a new benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests, combining synchronized RGB imagery and LiDAR from ground traversals with high‑resolution aerial imagery and multi‑altitude LiDAR over 370 hectares. The dataset includes accurate geo‑referenced 6‑DoF poses and spans 36 km of forest traversals, enabling systematic evaluation of visual, LiDAR, cross‑modal, and multi‑modal methods. Baseline experiments show LiDAR outperforms vision‑only approaches under severe viewpoint changes, while current multi‑modal fusion offers limited gains due to poor cross‑modal alignment, highlighting challenges in cross‑platform localisation and domain gaps.
Visual localization becomes extremely challenging in planetary-like terrains characterized by low texture, perceptual aliasing, harsh illumination, and sparse, weakly overlapping viewpoints induced by forward rover motion and unconstrained driving directions. Under these conditions, state-of-the-art image-to-image and image-to-map matching pipelines suffer significant performance degradation.
DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.
M3GD introduces a multimodal representation that fuses pre‑trained 2D image and 3D LiDAR foundation models for robotic novel view synthesis, avoiding the need for a separate cross‑modal translator. By projecting LiDAR onto the image latent grid and injecting the resulting geometry‑aware packets via a lightweight residual adapter, the method enhances both RGB and depth synthesis on the GrandTour dataset compared to an image‑only baseline. Ablation studies confirm that pixel‑aligned LiDAR content drives the performance gains, and real‑world deployment on a ground robot demonstrates a tunable quality–cost trade‑off.
Metric scale monocular geometry estimation has seen significant progress through large-scale data aggregation, yet current foundation models suffer from a persistent ''scale-collapse'' phenomenon: distant landmarks and vast landscapes are metrically underestimated. We hypothesize that this performance gap stems from a training data bottleneck, where existing metric-scale datasets are hardware-constrained to homogenous vehicle-captured LiDAR or short-range indoor scans, or consist of synthetic data that lacks the semantic complexity of the physical world.
M3GD is a novel approach for robotic novel view synthesis that fuses camera images and LiDAR point clouds without requiring a separate cross‑modal translator. By projecting LiDAR data onto the image latent grid and injecting it via a lightweight residual adapter, M3GD enhances both RGB and depth generation on the GrandTour dataset compared to image‑only baselines. Experiments on a ground robot confirm that the method can be deployed in real‑world scenarios with a tunable quality‑cost trade‑off.
GeoFF3D is a new feed‑forward 3D reconstruction method designed for large‑scale UAV mapping. It uses a coordinate‑anchored model that predicts camera poses and dense point maps directly in a gravity‑aligned Z‑up metric frame, while a spatial large‑scale reconstruction framework (SLRF) partitions images into overlapping chunks, propagates shared‑view priors, and aggregates local reconstructions hierarchically. Across nine aerial mapping blocks, GeoFF3D achieves the best average reconstruction quality, improving F@5 from 0.829 to 0.877, and can reconstruct 2,000 images in about five minutes.
arXiv:2506. 22784v2 Announce Type: replace-cross Abstract: Point-pixel registration between LiDAR point clouds and camera images is a fundamental yet challenging task in autonomous driving and robotic perception.
arXiv:2606. 20291v1 Announce Type: new Abstract: Remote sensing is increasingly relied upon to deliver actionable science for forest and wildfire risk management across large landscapes.