arXiv AI

Star-Fusion: A Multi-modal Transformer Architecture for Discrete Celestial Orientation via Spherical Topology

arXiv Machine Learning
Sep 24

PBLH Estimation from Satellite Radiances via a Dual-Encoder Transformer

The paper presents a dual‑encoder Transformer model for estimating Planetary Boundary Layer Height (PBLH) from satellite radiances, addressing challenges of multimodal, spatially incomplete data. It benchmarks eight different approaches, analyzes model reliance via grouped Shapley decomposition, and demonstrates that the proposed architecture achieves a mean absolute error of 155.8 m on a global test set, outperforming all baselines. On out‑of‑distribution data from the TEAMx campaign, the model attains 165.3 m MAE, better than a pixel‑wise baseline trained on the same data.

By Lorenzo Innocenti, Luca Catalano, Edoardo Arnaudo, Claudio Rossi, Salvatore Larosa, Domenico Cimini, Paolo Garza
arXiv Machine Learning
Sep 11

Bidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models

The paper introduces SolCloudLLM, a large language model–based framework that fuses sky‑image patches with time‑series data through bidirectional multimodal fusion for short‑term solar forecasting. Experiments on the SIRTA and SKIPP'D datasets show that SolCloudLLM outperforms existing baselines, achieving up to a 25.4% reduction in mean squared error, especially under cloudy conditions and in few‑shot scenarios.

By Ken Chen, Maneesha Perera, Wei Wang, Sachith Seneviratne, Hansani Weeratunge, Saman Halgamuge
Hugging Face Trending Papers
Sep 8

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.

arXiv Computer Vision
Sep 4

STARS-GS: Structure-Aware Regularized Gaussian Splatting for Large-Scale Aerial Surface Reconstruction

STARS-GS is a new structure‑aware 3D Gaussian Splatting framework designed for large‑scale aerial surface reconstruction. It introduces a scene partitioning strategy that preserves continuous scene elements, a neighborhood‑aware Gaussian organization that extends geometric constraints to local neighborhoods, and an adaptive surface regularization that tailors regularization strength to local geometry. Experiments on aerial photogrammetry benchmarks show that STARS‑GS improves the average F1‑score from 0.640 to 0.698, a relative gain of about 9.1%.

By Bocheng Li, Wenjuan Zhang, Jie Pan. Dongxu Han, Xuesong Ma, Yiling Yao, Yaning Wang