Benchmarking Generative Models for Weather Data Assimilation on Real Station Observations
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
SimCast‑S2S is a generative latent‑diffusion model designed for probabilistic subseasonal‑to‑seasonal precipitation forecasting. It tackles three key challenges: it uses a diffusion pipeline to capture uncertainty, operates in a compact latent space to enable efficient large‑ensemble generation, and leverages transfer learning with low‑rank adaptation to train on limited reanalysis data after pretraining on climate simulations. The model outperforms deep‑learning baselines and competes with, or surpasses, operational systems such as the ECMWF‑S2S baseline without requiring extensive post‑processing.
SimCast‑S2S is a generative latent‑diffusion framework designed for probabilistic subseasonal‑to‑seasonal precipitation forecasting. It tackles three key challenges: it uses a diffusion‑based generative pipeline for uncertainty quantification, operates in a compact latent space learned by VAEs for efficient large‑ensemble generation, and employs transfer learning with LoRA to overcome limited training data. On reanalysis data, it outperforms deep‑learning baselines and competes with or surpasses state‑of‑the‑art operational systems such as ECMWF‑S2S.
The paper introduces prequential posteriors, a Bayesian approach that uses a predictive‑sequential loss function to update deep generative forecasting models (DGFMs) when new data arrive. By adopting a consistency notion suitable for model misspecification, the authors prove that both the loss minimizer and the posterior concentrate on parameters with optimal predictive performance. Scalable inference is achieved with parallelisable waste‑free sequential Monte Carlo samplers that employ preconditioned gradient kernels, and the method is validated on synthetic and real meteorological time‑series data.
The paper introduces STORM, a one‑stage generative AI framework that reformulates Earth system data assimilation as diffusion‑based Bayesian posterior sampling, replacing costly PDE ensemble forecasts with scalable AI inference. STORM employs a spatiotemporal transformer with a global‑attention algorithm that reduces computational complexity from quadratic to linear, enabling high‑resolution, long‑context modeling. The system scales to 74,400 GPUs on Frontier, achieving 96–99 % strong‑scaling efficiency and up to 6 ExaFLOPs sustained BF16 throughput, while supporting 32,768‑member ensembles for uncertainty quantification in just 34 seconds on 4,096 GPUs, and demonstrates improved hurricane tracking and climate reanalysis accuracy.
arXiv:2609.39626v1 Announce Type: cross Abstract: Ensemble smoothers are the most successful and efficient techniques currently available for history matching. However, because these methods rely on...
arXiv:2605. 14285v2 Announce Type: replace-cross Abstract: Data assimilation (DA) estimates the state of an evolving dynamical system from noisy, partial observations, and is widely used in scientific simulation as well as weather and climate science.