arXiv:2602. 16579v2 Announce Type: replace-cross Abstract: Reliable global streamflow forecasting is essential for flood preparedness and water resource management, yet data-driven models often suffer from a performance gap when transitioning from historical reanalysis to operational forecast products.
By Maria Luisa Taccari, Kenza Tazi, Ois\'in M. Morrison, Andreas Grafberger, Juan Colonese, Corentin Carton de Wiart, Christel Prudhomme, Cinzia Mazzetti, Matthew Chantry, Florian Pappenberger
arXiv:2608.30976v1 Announce Type: new
Abstract: Real-world time-series forecasting is rarely a one-shot model invocation: practitioners must formulate tasks, connect data and models, incorporate doma...
By Xiaoyu Tao, Mingyue Cheng, Ze Guo, Bokai Pan, Qi Liu, Shijin Wang, Enhong Chen
Real-world time-series forecasting is rarely a one-shot model invocation: practitioners must formulate tasks, connect data and models, incorporate domain expertise, assess prediction plausibility, and...
arXiv:2607. 24588v1 Announce Type: new Abstract: Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events.
By Hang Ni, Weijia Zhang, Fan Liu, Mengqian Lu, Hao Liu
arXiv:2609.22702v1 Announce Type: new
Abstract: Accurate flood forecasts several days in advance are essential for flood control, water-resource management, and emergency response. Producing them at...
By Hong Zhang, John K. Hutchison, Rao Kotamarthi, Jeremy Feinstein, Haiwen Guan, Romit Maulik, Vijay P. Ramalingam, Jason Stock, Tom Wall
AFDBench is a new benchmark that evaluates how well large language models can generate professional Area Forecast Discussions (AFDs) for the National Weather Service by reasoning through structured AI weather forecast data. It contains 7,732 expert-written discussions paired with real forecast inputs and introduces three metrics—Met-Align, Style-Align, and Input-Grounding—to assess numerical accuracy, professional dialect adherence, and fidelity to source data. Zero-shot tests show open-source LLMs perform poorly on style and grounding, but reinforcement learning with Group Relative Policy Optimization nearly doubles style alignment and improves grounding, enabling a 7B-parameter model to write like a professional meteorologist.
By Manmeet Singh, Somnath Luitel, Prabhjot Singh, Manraaj Banga, Naveen Sudharsan, Josh Durkee