arXiv Computer Vision

Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning

Hugging Face Trending Papers
Sep 10

Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image Generators

The paper presents a method for enhancing coarse 5 m digital surface models (DSMs) to 0.5 m resolution by guiding a denoising diffusion process with high‑resolution spectral images. This approach transfers fine visual details—such as crisp outlines and roof structures—from the imagery into the elevation maps, yielding more accurate surface geometry than traditional interpolation or filtering. Experiments on Central European cities show that the resulting DSMs exhibit improved structural detail and overall quality.

arXiv Computer Vision
Sep 11

Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image Generators

The paper proposes a method to enhance coarse 5 m digital surface models (DSMs) to 0.5 m resolution by guiding the super‑resolution process with high‑resolution spectral images. It uses denoising diffusion to transfer image‑visible details, such as crisp outlines and roof structures, into the elevation maps, achieving more accurate surface geometry than traditional interpolation or filtering. Experiments on Central European cities show that the approach yields high‑quality DSMs with improved structural detail.

By Armand Mihai Nicolicioiu, Dominik Narnhofer, Nando Metzger, Daniel Panangian, Ksenia Bittner, Konrad Schindler
arXiv AI
Jun 29

Mind the Gap: Quantifying the Domain Gap in Cross-Sensor Diffusion Super-Resolution

arXiv:2606. 28039v1 Announce Type: cross Abstract: Demand for high-resolution satellite imagery has increased interest in super-resolution (SR) to bridge the spatial resolution gap between freely available missions such as Sentinel-2 and commercial systems like PlanetScope.

By Dawid Kope\'c, Katarzyna Jab{\l}o\'nska, Wojciech Koz{\l}owski, Maciej Zi\k{e}ba
arXiv Computer Vision
Sep 25

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

M3GD introduces a multimodal representation that fuses pre‑trained 2D image and 3D LiDAR foundation models for robotic novel view synthesis, avoiding the need for a separate cross‑modal translator. By projecting LiDAR onto the image latent grid and injecting the resulting geometry‑aware packets via a lightweight residual adapter, the method enhances both RGB and depth synthesis on the GrandTour dataset compared to an image‑only baseline. Ablation studies confirm that pixel‑aligned LiDAR content drives the performance gains, and real‑world deployment on a ground robot demonstrates a tunable quality–cost trade‑off.

By Yang Zhou, Jiuhong Xiao, Shizhao Ye, Long Quang, Carlos Nieto-Granda, Giuseppe Loianno