arXiv Machine Learning By Yixuan Jia, Siyi Chen, Yida Pan, Xiao Li, Lianghe Shi, Chanyong Jung, Haijie Yuan, Ismail Alkhouri, Yue Cynthia Wu, Saiprasad Ravishankar, Jeffrey A Fessler, Qing Qu

ForcingDAS: Unified and Robust Data Assimilation via Diffusion Forcing

Read the original on arXiv Machine Learning →

arXiv:2605. 14285v2 Announce Type: replace-cross Abstract: Data assimilation (DA) estimates the state of an evolving dynamical system from noisy, partial observations, and is widely used in scientific simulation as well as weather and climate science.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 2

Global Attention with Linear Complexity for Exascale Generative Data Assimilation in Earth System Prediction

The paper introduces STORM, a one‑stage generative AI framework that reformulates Earth system data assimilation as diffusion‑based Bayesian posterior sampling, replacing costly PDE ensemble forecasts with scalable AI inference. STORM employs a spatiotemporal transformer with a global‑attention algorithm that reduces computational complexity from quadratic to linear, enabling high‑resolution, long‑context modeling. The system scales to 74,400 GPUs on Frontier, achieving 96–99 % strong‑scaling efficiency and up to 6 ExaFLOPs sustained BF16 throughput, while supporting 32,768‑member ensembles for uncertainty quantification in just 34 seconds on 4,096 GPUs, and demonstrates improved hurricane tracking and climate reanalysis accuracy.

By Xiao Wang, Zezhong Zhang, Isaac Lyngaas, Hong-Jun Yoon, Jong-Youl Choi, Siming Liang, Janet Wang, Hristo G. Chipilski, Ashwin M. Aji, Feng Bao, Peter Jan van Leeuwen, Dan Lu, Guannan Zhang
arXiv AI
2d ago

Benchmarking Generative Models for Weather Data Assimilation on Real Station Observations

This study introduces the first controlled benchmark of generative models for weather data assimilation using real station observations from 11,849 NOAA MADIS stations across the U.S. It evaluates key design choices—diffusion vs. flow matching, pixel vs. latent-space formulations, and inference-time conditioning strategies—against a classical 3D-Var baseline. The benchmark finds that learned generative priors and full-gradient guidance improve RMSE over ERA5, while other design variations offer minimal benefit, especially under sparse observation conditions.

By Ruizhe Huang, Qidong Yang, Jonathan Giezendanner, Sherrie Wang