arXiv Machine Learning

Prequential posteriors

The paper introduces prequential posteriors, a Bayesian approach that uses a predictive‑sequential loss function to update deep generative forecasting models (DGFMs) when new data arrive. By adopting a consistency notion suitable for model misspecification, the authors prove that both the loss minimizer and the posterior concentrate on parameters with optimal predictive performance. Scalable inference is achieved with parallelisable waste‑free sequential Monte Carlo samplers that employ preconditioned gradient kernels, and the method is validated on synthetic and real meteorological time‑series data.

arXiv Machine Learning
Aug 28

SimCast-S2S: An Efficient Generative Model for Subseasonal Precipitation Forecasting via Transfer Learning from Climate Simulations

SimCast‑S2S is a generative latent‑diffusion model designed for probabilistic subseasonal‑to‑seasonal precipitation forecasting. It tackles three key challenges: it uses a diffusion pipeline to capture uncertainty, operates in a compact latent space to enable efficient large‑ensemble generation, and leverages transfer learning with low‑rank adaptation to train on limited reanalysis data after pretraining on climate simulations. The model outperforms deep‑learning baselines and competes with, or surpasses, operational systems such as the ECMWF‑S2S baseline without requiring extensive post‑processing.

By Hiep V. Dang, Antonios Mamalakis
arXiv Machine Learning
Sep 4

SimCast-S2S: A Computationally Efficient Diffusion Model for Subseasonal Precipitation Forecasting

SimCast‑S2S is a generative latent‑diffusion framework designed for probabilistic subseasonal‑to‑seasonal precipitation forecasting. It tackles three key challenges: it uses a diffusion‑based generative pipeline for uncertainty quantification, operates in a compact latent space learned by VAEs for efficient large‑ensemble generation, and employs transfer learning with LoRA to overcome limited training data. On reanalysis data, it outperforms deep‑learning baselines and competes with or surpasses state‑of‑the‑art operational systems such as ECMWF‑S2S.

By Hiep V. Dang, Antonios Mamalakis
arXiv Machine Learning
Aug 7

Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulations

arXiv:2608. 06107v1 Announce Type: new Abstract: Machine learning offers a promising avenue to accelerate physical simulations by replacing computationally expensive traditional Partial Differential Equation (PDE) solvers with fast, differentiable surrogate models.

By Guillaume Couairon, Alexis Jacq, Yu-Han Wu, Renu Singh, Yana Hasson, Quentin Berthet, Romuald Elie
arXiv AI
Sep 2

Global Attention with Linear Complexity for Exascale Generative Data Assimilation in Earth System Prediction

The paper introduces STORM, a one‑stage generative AI framework that reformulates Earth system data assimilation as diffusion‑based Bayesian posterior sampling, replacing costly PDE ensemble forecasts with scalable AI inference. STORM employs a spatiotemporal transformer with a global‑attention algorithm that reduces computational complexity from quadratic to linear, enabling high‑resolution, long‑context modeling. The system scales to 74,400 GPUs on Frontier, achieving 96–99 % strong‑scaling efficiency and up to 6 ExaFLOPs sustained BF16 throughput, while supporting 32,768‑member ensembles for uncertainty quantification in just 34 seconds on 4,096 GPUs, and demonstrates improved hurricane tracking and climate reanalysis accuracy.

By Xiao Wang, Zezhong Zhang, Isaac Lyngaas, Hong-Jun Yoon, Jong-Youl Choi, Siming Liang, Janet Wang, Hristo G. Chipilski, Ashwin M. Aji, Feng Bao, Peter Jan van Leeuwen, Dan Lu, Guannan Zhang