arXiv Machine Learning

Calibrating subgrid parametrizations of single-column ocean models via simulation-based inference

arXiv Machine Learning
Jun 11

Deep Learning of Solver-Aware Turbulence Closures from Nudged LES Dynamics

arXiv:2604. 23874v3 Announce Type: replace-cross Abstract: The differentiable physics paradigm may be leveraged as an a-posteriori approach for discovering turbulence closure models by embedding a neural network parameterization directly inside the solver and optimizing it given potentially sparse target data.

By Ashwin Suriyanarayanan, Dibyajyoti Chakraborty, Romit Maulik
arXiv Machine Learning
6d ago

HClimRep-Ocean: A Global Ocean Emulator on an Unstructured Mesh

HClimRep‑Ocean is a machine‑learning emulator that operates directly on the native unstructured mesh of the FESOM2 ocean model, trained on a 209‑year AWI‑CM3 control run and run without atmospheric forcing except at initialization. It shows strong skill for current forecasts at 30‑day lead times, outperforming all references, while temperature and salinity forecasts are best served by a damped‑anomaly persistence approach. In independent OceanBench testing, a reanalysis‑trained variant achieves the lowest RMSE against GLORYS reanalysis, demonstrating the competitiveness of the native‑mesh approach.

By Kacper Nowak, Aleksei Koldunov, Nikolay Koldunov, Savvas Melidonis, Ankit Patnala, Simon Grasse, Julius Polz, Christian Lessig, Martin Schultz, Thomas Jung
arXiv AI
Jul 22

Incomplete Observations Boost Evolutionary Performance in Ocean Modeling

arXiv:2607. 19147v1 Announce Type: cross Abstract: Data-driven methods have revolutionized ocean modeling, yet current approaches rely heavily on complete reanalysis datasets, imposing computational constraints and limiting model performance to that of the training data.

By Yangyang Kong, Yutong Jiang, Yanhai Gan, Junyu Dong, Feng Gao, Xiaopei Lin
arXiv Statistics ML
Sep 3

The Ensemble Kalman Inversion Race

The paper compares different Ensemble Kalman methods for calibrating climate model parameters by minimizing the misfit between modeled and observed climate statistics. It conducts systematic numerical experiments on Lorenz-type models, including neural network parameterizations, to evaluate computational efficiency and accuracy of each method. The study examines how prior information and dimensionality affect the cost of these methods.

By Rebecca Gjini, Matthias Morzfeld, Oliver R. A. Dunbar, Tapio Schneider
arXiv AI
Aug 18

OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations

OceanDepths is the first open, global, AI‑ready dataset that pairs satellite‑derived sea surface temperature, salinity, and height with co‑located EN4 subsurface temperature and salinity profiles, complemented by GLORYS12 reanalysis data. It covers 2000–2024 at 0.1°×0.1° spatial resolution and weekly temporal resolution, providing over 9.5 million paired profiles interpolated to 50 depth levels. The dataset’s 4‑D multivariate structure, high resolution, long temporal extent, and extreme sparsity of subsurface observations make it a challenging testbed for novel AI methods, with demonstrated use in subsurface state reconstruction and potential for observation‑based forecasting.

By Simon Donike, Ruben Cartuyvels, Antonino Ian Ferola, Elisa Carli, Diego Fernandez Prieto, Marie-Helene Rio
arXiv AI
Jun 26

Sampling sea state using a diffusion model

arXiv:2606. 26389v1 Announce Type: cross Abstract: Sea state prediction is essential for operational maritime applications and coupled earth system modeling, yet current spectral wave models remain computationally prohibitive for many use cases, including online coupling to climate simulations and making probabilistic (ensemble-based) predictions.

By Jiarong Wu, Bertrand Chapron, Laure Zanna
arXiv Machine Learning
Sep 3

Source Distribution Estimation by Posterior Averaging

The paper introduces a new approach to source distribution estimation (SDE) in simulation-based science, addressing limitations of existing methods that rely on a fixed surrogate likelihood. By employing an expectation‑maximization framework, the authors iteratively train an amortized posterior on fresh simulations (E‑step) and refit the source distribution to the posterior’s average (M‑step). Two parameterizations are explored: separate source and posterior flows, and a single shared conditional flow, with experiments on three benchmark tasks showing improved performance over fixed surrogate and iterated baseline methods, notably achieving higher data‑space C2ST scores on the Lotka–Volterra benchmark.

By Trung-Dung Hoang, Lisa M. Koch