arXiv AI

SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

arXiv:2607. 24588v1 Announce Type: new Abstract: Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events.

arXiv AI
Sep 11

Agent-Based ML-LLM Fusion with Self-Optimizing Prompts for Plateau Weather Alerts

The paper introduces SmartWeatherAgent, a three‑stage architecture that combines intent recognition, hazard prediction, and reasoning‑enhanced generation to improve tourism meteorological services. It fuses rule‑based methods with large language models and a LightGBM model enriched with highland‑specific features, achieving an F1‑Macro score of 0.605 and 1.60 ms latency on high‑wind, precipitation, and low‑temperature events. A 12‑round micro‑step prompt self‑optimization loop raises the composite warning quality score from 4.2 to 8.9, with notable gains in data source citation, physical mechanism explanation, and scientific rigor through explicit uncertainty statements.

By Shuai Yan, Yang Xu, Shan He
arXiv AI
Jul 7

HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis

arXiv:2603. 01121v2 Announce Type: replace Abstract: While deep learning-based weather forecasting paradigms have made significant strides, addressing extreme weather diagnostics remains a formidable challenge.

By Shuo Tang, Jiadong Zhang, Gengxian Zhou, Qizhao Jin, Qinxuan Wang, Yi Hu, Ning Hu, Hongchang Ren, Lingli He, Shiming Xiang, Jingtao Ding, Jian Xu, Jiaolan Fu, Cheng-Lin Liu
arXiv Machine Learning
Jul 28

HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows

arXiv:2607. 23983v1 Announce Type: cross Abstract: Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer.

By Qingyi Yang, Siqian Qiu, Bing Li, Xu Shan, Jia Feng, Shunan Zhou, Xudong Zhou, Tiantian Xing, Jiale Guo, Xiaoyi Dong, Gaoyu Liu, Xiaohuan Liu, Haiqing Pu, Qingwen Deng, Xun Zhang, Zhongrun Xiang, Haiyang Qian, Ying Yan, Yongkang Xu, Nuo Lei, Tianlong Jia, Baoying Shan, Carlo De Michele
arXiv Machine Learning
Aug 27

AFDBench: A Reasoning-First AI Scientist for NationalWeather Service Forecast Discussions

AFDBench is a new benchmark that evaluates how well large language models can generate professional Area Forecast Discussions (AFDs) for the National Weather Service by reasoning through structured AI weather forecast data. It contains 7,732 expert-written discussions paired with real forecast inputs and introduces three metrics—Met-Align, Style-Align, and Input-Grounding—to assess numerical accuracy, professional dialect adherence, and fidelity to source data. Zero-shot tests show open-source LLMs perform poorly on style and grounding, but reinforcement learning with Group Relative Policy Optimization nearly doubles style alignment and improves grounding, enabling a 7B-parameter model to write like a professional meteorologist.

By Manmeet Singh, Somnath Luitel, Prabhjot Singh, Manraaj Banga, Naveen Sudharsan, Josh Durkee
arXiv AI
Aug 25

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

arXiv:2608.23058v1 Announce Type: new Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external too...

By Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu, Wei Wei, Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng
arXiv AI
Aug 6

Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals

arXiv:2607. 21597v2 Announce Type: replace Abstract: Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event prediction accuracy, not the operational coherence of a continuous risk signal.

By Nicolas Caron, Christophe Guyeux, Hassan Noura, Maxime Coulmeau, Benjamin Aynes