An adaptive and evolvable deep reinforcement learning framework for weather prediction
arXiv:2608. 09948v1 Announce Type: cross Abstract: No single AI weather model excels at all variables, pressure levels, and lead times.
AFDBench is a new benchmark that evaluates how well large language models can generate professional Area Forecast Discussions (AFDs) for the National Weather Service by reasoning through structured AI weather forecast data. It contains 7,732 expert-written discussions paired with real forecast inputs and introduces three metrics—Met-Align, Style-Align, and Input-Grounding—to assess numerical accuracy, professional dialect adherence, and fidelity to source data. Zero-shot tests show open-source LLMs perform poorly on style and grounding, but reinforcement learning with Group Relative Policy Optimization nearly doubles style alignment and improves grounding, enabling a 7B-parameter model to write like a professional meteorologist.
arXiv:2608. 09948v1 Announce Type: cross Abstract: No single AI weather model excels at all variables, pressure levels, and lead times.
arXiv:2607. 23983v1 Announce Type: cross Abstract: Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer.
arXiv:2606. 26421v1 Announce Type: new Abstract: State-of-the-art medium-range AI weather models can outperform traditional Numerical Weather Prediction (NWP) but require massive training budgets.
The paper introduces SmartWeatherAgent, a three‑stage architecture that combines intent recognition, hazard prediction, and reasoning‑enhanced generation to improve tourism meteorological services. It fuses rule‑based methods with large language models and a LightGBM model enriched with highland‑specific features, achieving an F1‑Macro score of 0.605 and 1.60 ms latency on high‑wind, precipitation, and low‑temperature events. A 12‑round micro‑step prompt self‑optimization loop raises the composite warning quality score from 4.2 to 8.9, with notable gains in data source citation, physical mechanism explanation, and scientific rigor through explicit uncertainty statements.
arXiv:2606. 02663v1 Announce Type: cross Abstract: Recent advances in machine learning have produced probabilistic weather forecasting models comparable to state-of-the-art numerical weather predictors.
The paper introduces Probabilistic Bias Correction (PBC), a machine learning framework that learns to correct historical probabilistic forecasts, thereby reducing systematic errors in subseasonal weather predictions. Applied to leading dynamical and AI models from ECMWF, PBC doubles the AI system’s modest subseasonal skill and improves the operationally-debiased dynamical model for most pressure, temperature, and precipitation targets. In ECMWF’s 2025 real‑time forecasting competition, PBC’s global forecasts ranked first across all weather variables and lead times, outperforming multiple operational and ensemble models.
arXiv:2604. 16238v2 Announce Type: replace Abstract: Decision-makers rely on weather forecasts to plant crops, manage wildfires, allocate water and energy, and prepare for weather extremes.
State-of-the-art medium-range AI weather models can outperform traditional Numerical Weather Prediction (NWP) but require massive training budgets. This restricts usage for under-resourced groups and severely limits fast model iteration.
arXiv:2604. 12306v3 Announce Type: replace-cross Abstract: Climate decision-making in the GCC states increasingly demands systems that can translate heterogeneous scientific and policy evidence into actionable guidance, yet general-purpose large language models (LLMs) remain weak both in region-specific climate knowledge and grounded interaction with geospatial and forecasting tools.
arXiv:2607. 24588v1 Announce Type: new Abstract: Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events.
The paper presents a graph-transformer AI weather model that is fine‑tuned with high‑resolution IMERG precipitation observations, moving beyond the traditional reliance on the ERA5 reanalysis dataset. This approach yields up to a 19% improvement in medium‑range continuous ranked probability scores and a 57% better Brier skill score for extreme rainfall compared to leading operational models, while also excelling in tropical storm and drizzle prediction. The study demonstrates that directly incorporating observation‑based precipitation data into AI training can markedly enhance forecast accuracy, though physics‑based models still outperform for the heaviest events.
To address insufficient contextualization, weak generalization, and poor scenario adaptation in tourism meteorological services, we propose SmartWeatherAgent--a unified three-stage architecture integr...