The paper introduces SmartWeatherAgent, a three‑stage architecture that combines intent recognition, hazard prediction, and reasoning‑enhanced generation to improve tourism meteorological services. It fuses rule‑based methods with large language models and a LightGBM model enriched with highland‑specific features, achieving an F1‑Macro score of 0.605 and 1.60 ms latency on high‑wind, precipitation, and low‑temperature events. A 12‑round micro‑step prompt self‑optimization loop raises the composite warning quality score from 4.2 to 8.9, with notable gains in data source citation, physical mechanism explanation, and scientific rigor through explicit uncertainty statements.
By Shuai Yan, Yang Xu, Shan He
arXiv:2607. 24588v1 Announce Type: new Abstract: Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events.
By Hang Ni, Weijia Zhang, Fan Liu, Mengqian Lu, Hao Liu
arXiv:2603. 01121v2 Announce Type: replace Abstract: While deep learning-based weather forecasting paradigms have made significant strides, addressing extreme weather diagnostics remains a formidable challenge.
By Shuo Tang, Jiadong Zhang, Gengxian Zhou, Qizhao Jin, Qinxuan Wang, Yi Hu, Ning Hu, Hongchang Ren, Lingli He, Shiming Xiang, Jingtao Ding, Jian Xu, Jiaolan Fu, Cheng-Lin Liu
arXiv:2607. 23983v1 Announce Type: cross Abstract: Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer.
By Qingyi Yang, Siqian Qiu, Bing Li, Xu Shan, Jia Feng, Shunan Zhou, Xudong Zhou, Tiantian Xing, Jiale Guo, Xiaoyi Dong, Gaoyu Liu, Xiaohuan Liu, Haiqing Pu, Qingwen Deng, Xun Zhang, Zhongrun Xiang, Haiyang Qian, Ying Yan, Yongkang Xu, Nuo Lei, Tianlong Jia, Baoying Shan, Carlo De Michele
AFDBench is a new benchmark that evaluates how well large language models can generate professional Area Forecast Discussions (AFDs) for the National Weather Service by reasoning through structured AI weather forecast data. It contains 7,732 expert-written discussions paired with real forecast inputs and introduces three metrics—Met-Align, Style-Align, and Input-Grounding—to assess numerical accuracy, professional dialect adherence, and fidelity to source data. Zero-shot tests show open-source LLMs perform poorly on style and grounding, but reinforcement learning with Group Relative Policy Optimization nearly doubles style alignment and improves grounding, enabling a 7B-parameter model to write like a professional meteorologist.
By Manmeet Singh, Somnath Luitel, Prabhjot Singh, Manraaj Banga, Naveen Sudharsan, Josh Durkee
arXiv:2511.20109v2 Announce Type: replace
Abstract: Climate science demands automated workflows to transform comprehensive questions into data-driven statements across massive, heterogeneous datasets...
By Chenyue Li, Hyeonjae Kim, Wen Deng, Mengxi Jin, Wen Huang, Mengqian Lu, Binhang Yuan
arXiv:2608.23058v1 Announce Type: new
Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external too...
By Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu, Wei Wei, Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng
KairosAgent is an agentic framework that combines a large language model (LLM) reasoner with a time series foundation model (TSFM) forecaster to tackle cross‑domain multimodal time series forecasting. It dynamically invokes analytical tools to improve the LLM’s numerical comprehension and semantic reasoning, then fuses the reasoning outcomes into the TSFM pipeline for more accurate predictions. The approach is further enhanced by a curated large‑scale trajectory corpus and a reinforcement learning paradigm with multi‑turn refinement and turn‑level credit assignment, achieving superior zero‑shot forecasting performance.
By Kun Feng, Ziwei Shan, Yuchen Fang, Yiyang Tan, Sihan Lu, Shuqi Gu, Xingyu Lu, Lintao Ma, Kan Ren
arXiv:2608. 11679v1 Announce Type: new Abstract: Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems.
By Touseef Hasan, Mounika Ghanta, Souvika Sarkar, Ujjwal Guin
The study presents a deployment‑aware framework for forecasting spring discharge and groundwater levels in the Edwards Aquifer over 1‑12 week horizons using 79 years of hydroclimatic data. Five machine‑learning families—extreme gradient boosting, extremely randomized trees, LSTM, CNN, and Transformers—were compared, with extreme gradient boosting consistently delivering the highest reliability (R² ≥ 0.94) and strong agreement with operational drought thresholds. The validated models were integrated into a five‑agent operational architecture that automates data acquisition, model selection, prediction, threshold monitoring, verification, literature retrieval, and reporting.
By Pramod Lekhak, Chetan Sharma, Hakan Ba\c{s}a\u{g}ao\u{g}lu, F. Paul Bertetti, Debaditya Chakraborty
arXiv:2604. 12306v3 Announce Type: replace-cross Abstract: Climate decision-making in the GCC states increasingly demands systems that can translate heterogeneous scientific and policy evidence into actionable guidance, yet general-purpose large language models (LLMs) remain weak both in region-specific climate knowledge and grounded interaction with geospatial and forecasting tools.
By Muhammad Umer Sheikh, Khawar Shehzad, Salman Khan, Fahad Shahbaz Khan, Muhammad Haris Khan
The paper introduces a two‑stage training framework that combines Supervised Fine‑Tuning (SFT) and Direct Preference Optimization (DPO) to improve multimodal disaster severity assessment. It creates two datasets—ReasoningSet for validated rationales and PreferenceSet for paired rationales—using a single Human‑in‑the‑Loop workflow. Experiments on InternVL‑3‑8B and LLaVA‑1.5‑7B show that SFT boosts classification accuracy and Macro‑F1, while DPO further enhances interpretability and alignment with human judgment.
By Yuanjun Zhang, Fuzel Ahamed Shaik, Suvojit Acharjee, Fahad Khalid, Mourad Oussalah