arXiv Machine Learning

Replacing Tunable Parameters in Weather and Climate Models with State-Dependent Functions using Reinforcement Learning

arXiv:2601. 04268v3 Announce Type: replace Abstract: Weather and climate models rely on parametrisations to represent unresolved sub-grid processes.

arXiv Machine Learning
Sep 3

Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling

The study couples the Met Office Unified Model with distributed reinforcement learning agents, using a DDPG actor that applies bounded potential‑temperature corrections across 70 vertical levels. Training is performed on ten nudged forecasts, after which the frozen policy is evaluated in a non‑nudged forecast, demonstrating numerical stability. The learned policy reduces Z₅₀₀ MAE in four of six latitude bands—up to 45.8% in the northern tropics—and decreases MSLP error by up to 27.3% in certain bands, indicating promising bias‑correction potential.

By Pritthijit Nath, Sebastian Schemm, Peter Haynes, Emily Shuckburgh, Mark Webb
Hugging Face Trending Papers
Sep 2

Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling

The study couples the Met Office Unified Model with distributed reinforcement learning agents that apply bounded temperature corrections across 70 vertical levels. Training on ten nudged forecasts and evaluating on a non‑nudged run, the learned policy remains numerically stable and improves forecast accuracy, reducing Z₅₀₀ MAE by up to 45.8% in tropical bands and MSLP error by up to 27.3% in certain latitudes. This experiment demonstrates the feasibility of online RL for bias correction in operational weather models.

arXiv Machine Learning
Sep 21

Learning to Advect: A Neural Semi-Lagrangian Architecture for Weather Forecasting

arXiv:2601.21151v3 Announce Type: replace Abstract: Machine-learning approaches to weather forecasting often employ a monolithic architecture in which distinct physical mechanisms, such as advection,...

By Carlos A. Pereira, St\'ephane Gaudreault, Valentin Dallerit, Christopher Subich, Shoyon Panday, Siqi Wei, Sasa Zhang, Siddharth Rout, Eldad Haber, Raymond J. Spiteri, David Millard
arXiv Machine Learning
Aug 19

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.

By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
Hugging Face Trending Papers
Jun 17

Optimal scenario design for climate emulation

As deep learning for physical systems continues to grow in popularity, efforts to improve generalizability have primarily focused on designing architectures that embed physical constraints. However, for machine-learning surrogate climate models (emulators), we show that the low structural diversity in existing scenarios commonly used to generate training data places a ceiling on predictive skill.

arXiv AI
Sep 10

Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems

The paper introduces an action‑conditioned world‑modeling framework that turns Earth‑system simulator trajectories into training data for controllable state‑transition learning. By pretraining on naturally observed state changes as implicit action supervision and using masked response learning, the model can infer unobserved variables and learn coupled system dependencies. Experiments on ecosystem dynamics across six global regions demonstrate that the model maintains long‑horizon emulation accuracy while enabling structural interventions and coherent responses in coupled ecosystem‑cycle variables.

By Zhihao Wang, Ruichen Wang, Ruohan Li, Lei Ma, George Hurtt, Xiaowei Jia, Gengchen Mai, Shaowen Wang, Yiqun Xie
arXiv Machine Learning
Jul 20

Comparative Field Deployment of Reinforcement Learning and Model Predictive Control for Residential HVAC

arXiv:2510. 01475v2 Announce Type: replace-cross Abstract: Model Predictive Control (MPC) has demonstrated significant performance improvements over today's control methods for residential Heating, Ventilation, and Air Conditioning (HVAC), but deploying MPC often requires substantial engineering effort.

By Ozan Baris Mulayim, Elias N. Pergantis, Levi D. Reyes Premer, Bingqing Chen, Guannan Qu, Kevin J. Kircher, Mario Berg\'es
arXiv AI
Jun 30

Agile Reinforcement Learning through Separable Neural Architecture and Applications

arXiv:2601. 23225v2 Announce Type: replace-cross Abstract: Deep reinforcement learning (RL) is increasingly deployed in resource-constrained environments, yet go-to function approximators - multilayer perceptrons (MLPs) - are often parameter-inefficient due to an imperfect inductive bias for the smooth structure of many value functions.

By Rajib Mostakim, Reza T. Batley, Sourav Saha