The study couples the Met Office Unified Model with distributed reinforcement learning agents, using a DDPG actor that applies bounded potential‑temperature corrections across 70 vertical levels. Training is performed on ten nudged forecasts, after which the frozen policy is evaluated in a non‑nudged forecast, demonstrating numerical stability. The learned policy reduces Z₅₀₀ MAE in four of six latitude bands—up to 45.8% in the northern tropics—and decreases MSLP error by up to 27.3% in certain bands, indicating promising bias‑correction potential.
By Pritthijit Nath, Sebastian Schemm, Peter Haynes, Emily Shuckburgh, Mark Webb
The study couples the Met Office Unified Model with distributed reinforcement learning agents that apply bounded temperature corrections across 70 vertical levels. Training on ten nudged forecasts and evaluating on a non‑nudged run, the learned policy remains numerically stable and improves forecast accuracy, reducing Z₅₀₀ MAE by up to 45.8% in tropical bands and MSLP error by up to 27.3% in certain latitudes. This experiment demonstrates the feasibility of online RL for bias correction in operational weather models.
arXiv:2609.24882v1 Announce Type: new
Abstract: Hybrid AI-physics climate modeling aims to improve coarse (~100km-resolution) Earth system models by learning to parameterize subgrid processes from hi...
By Jurij Sch\"onfeld, Tom Beucler, Julien Savre, Steven Sherwood, Veronika Eyring
arXiv:2605.16929v2 Announce Type: replace
Abstract: Global climate models are essential tools to simulate past and potential future pathways of climate change, as well as associated climate impacts....
By Graham Clyne, Julia Kaltenborn, Peer Nowack, Claire Monteleoni, Anastase Charantonis
arXiv:2608. 09948v1 Announce Type: cross Abstract: No single AI weather model excels at all variables, pressure levels, and lead times.
By Qiang Wu, Han Li, Jianping Huang
arXiv:2601.21151v3 Announce Type: replace
Abstract: Machine-learning approaches to weather forecasting often employ a monolithic architecture in which distinct physical mechanisms, such as advection,...
By Carlos A. Pereira, St\'ephane Gaudreault, Valentin Dallerit, Christopher Subich, Shoyon Panday, Siqi Wei, Sasa Zhang, Siddharth Rout, Eldad Haber, Raymond J. Spiteri, David Millard
arXiv:2609.25505v1 Announce Type: cross
Abstract: Rapid intensification (RI) remains one of the most consequential and difficult aspects of tropical cyclone (TC) forecasting. Although full-physics nu...
By Shijie Xiao, Jonathan Lin, Thomas Ehrmann, Ali Sarhadi
Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.
By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
As deep learning for physical systems continues to grow in popularity, efforts to improve generalizability have primarily focused on designing architectures that embed physical constraints. However, for machine-learning surrogate climate models (emulators), we show that the low structural diversity in existing scenarios commonly used to generate training data places a ceiling on predictive skill.
The paper introduces an action‑conditioned world‑modeling framework that turns Earth‑system simulator trajectories into training data for controllable state‑transition learning. By pretraining on naturally observed state changes as implicit action supervision and using masked response learning, the model can infer unobserved variables and learn coupled system dependencies. Experiments on ecosystem dynamics across six global regions demonstrate that the model maintains long‑horizon emulation accuracy while enabling structural interventions and coherent responses in coupled ecosystem‑cycle variables.
By Zhihao Wang, Ruichen Wang, Ruohan Li, Lei Ma, George Hurtt, Xiaowei Jia, Gengchen Mai, Shaowen Wang, Yiqun Xie
arXiv:2510. 01475v2 Announce Type: replace-cross Abstract: Model Predictive Control (MPC) has demonstrated significant performance improvements over today's control methods for residential Heating, Ventilation, and Air Conditioning (HVAC), but deploying MPC often requires substantial engineering effort.
By Ozan Baris Mulayim, Elias N. Pergantis, Levi D. Reyes Premer, Bingqing Chen, Guannan Qu, Kevin J. Kircher, Mario Berg\'es
arXiv:2601. 23225v2 Announce Type: replace-cross Abstract: Deep reinforcement learning (RL) is increasingly deployed in resource-constrained environments, yet go-to function approximators - multilayer perceptrons (MLPs) - are often parameter-inefficient due to an imperfect inductive bias for the smooth structure of many value functions.
By Rajib Mostakim, Reza T. Batley, Sourav Saha