arXiv Machine Learning

Evaluating Large Language Models for Forced Outage Risk Prediction: Benefits and Comparison to Machine Learning

This study evaluates large language models (LLMs) for predicting weather‑related forced outage risk in a distribution grid using a zero‑shot approach without labeled training data. The task is framed as binary severity classification over 3h, 6h, and 12h horizons, leveraging six years of outage records and high‑resolution weather data from central Texas. Four zero‑shot LLMs are compared to two supervised classifiers under two input settings—current weather observations and forecast data—showing that supervised models lead on macro‑F1 and precision, while newer LLMs achieve competitive scores and offer complementary strengths in reasoning and geographic scalability.

arXiv AI
Sep 3

OutageDiT: A Generative Foundation Model for Power Outage Forecasting and Scenario Simulation

OutageDiT is a generative foundation model that produces seven‑day power‑outage trajectories at quarter‑hour resolution, trained on nationwide outage and weather data. It uses a condition encoder to process historical context and future covariates, and a shallow flow decoder to generate full trajectories, enabling point forecasting, uncertainty quantification, and conditional event simulation. The model outperforms strong baselines on forecasting benchmarks and can transfer zero‑shot to unseen regions, linking outage simulation to operational planning under uncertainty.

By Yunqin Zhu, Feng Qiu, Yao Xie
arXiv AI
Aug 25

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

arXiv:2608.23058v1 Announce Type: new Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external too...

By Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu, Wei Wei, Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng
arXiv Machine Learning
6d ago

Every Fixed Metric Has a Blind Spot: A Learned Atmospheric Critic for Scoring Forecast Realism

The paper introduces a learned atmospheric critic that discriminates between real weather data and model outputs to produce a realism score. Unlike fixed metrics, the discriminator adapts to the specific failure modes of a given model, effectively detecting various synthetic corruptions in ERA5 data. Experiments show the learned critic outperforms existing metrics and reveals that realism decreases with longer forecast lead times, favoring numerical over machine‑learning models.

By Younes Elberkennou, Dmitri Demler, Thierry Meier, Luca Rispoli, Fanny Lehmann, Joel Oskarsson
arXiv Machine Learning
Aug 19

Evaluating and improving crop-yield forecasting methods during extreme drought

The study evaluates crop‑yield forecasting methods for the 2012 Midwestern US drought, comparing non‑deep learning machine learning models with a deep learning model (VITA) using 16 meteorological predictors. It highlights challenges such as distributional dissimilarity between training and test data, spatial and temporal sparsity, and demonstrates that sample weighting and feature selection improve non‑deep learning models but not VITA. The work contrasts deep versus non‑deep learning approaches and shows how modifications can mitigate issues arising from extreme drought conditions.

By Shrey Gupta, Yi Ming, George Mohler
arXiv Machine Learning
Aug 7

Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction

arXiv:2608. 05265v1 Announce Type: new Abstract: Prediction of post-wildfire debris flows is critical for mitigating hazards to communities, infrastructure, and resources during intense rainfall in recently burned areas.

By Quinn Ledingham, Zhengsen Xu, Yimin Zhu, Zack Dewis, Mabel Heffring, Saeid Taleghanidoozdoozan, Motasem Alkayid, Megan Greenwood, Lincoln Linlin Xu
arXiv Machine Learning
Aug 20

Scalable Geospatial Machine Learning for Power-Line Asset Risk: Integrating Remote Sensing for Lightning and Vegetation Risk Modelling

The paper presents a modular, scalable framework for estimating the probability of failure (PoF) of power‑line assets using geospatial machine learning. It integrates diverse environmental predictors—topography, vegetation indices, lightning climatology, proximity features, and operational records—to model vegetation‑ and lightning‑related failure modes. The architecture is designed to be computationally efficient, easily extensible to new data sources, and suitable for large‑scale utility deployment, enabling asset‑level risk stratification for inspection and resilience planning.

By Artur Sokolovsky, Bhavik Merai, Moe Jafari, Muen Chen
arXiv Machine Learning
1d ago

WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction

WildfireSpreadBench evaluates machine‑learning models for predicting next‑day wildfire spread, comparing five discriminative and one generative architecture on the WildfireSpreadTS dataset. The study shows that model rankings differ markedly when using Average Precision versus threshold‑dependent metrics such as F1 and IoU, revealing three distinct prediction profiles—over‑predicting, balanced, and under‑predicting—that AP alone cannot distinguish. Expanding input channels modestly affected AP, underscoring that AP may favor models with predictions poorly suited for operational use.

By Arin Gopakumar, Marco Pannozzo
arXiv Machine Learning
Jun 9

Zero and Few Shot Load Forecasting with Large Language Models

arXiv:2411. 11350v2 Announce Type: replace Abstract: Deep learning models have shown strong performance in load forecasting, but they generally require large amounts of data for model training before being applied to new scenarios, which limits their effectiveness in data-scarce scenarios.

By Wenlong Liao, Chengrui Zhang, Zhe Yang, Mengshuo Jia, Christian Rehtanz, Jiannong Fang, Fernando Port\'e-Agel
arXiv Machine Learning
Aug 27

AFDBench: A Reasoning-First AI Scientist for NationalWeather Service Forecast Discussions

AFDBench is a new benchmark that evaluates how well large language models can generate professional Area Forecast Discussions (AFDs) for the National Weather Service by reasoning through structured AI weather forecast data. It contains 7,732 expert-written discussions paired with real forecast inputs and introduces three metrics—Met-Align, Style-Align, and Input-Grounding—to assess numerical accuracy, professional dialect adherence, and fidelity to source data. Zero-shot tests show open-source LLMs perform poorly on style and grounding, but reinforcement learning with Group Relative Policy Optimization nearly doubles style alignment and improves grounding, enabling a 7B-parameter model to write like a professional meteorologist.

By Manmeet Singh, Somnath Luitel, Prabhjot Singh, Manraaj Banga, Naveen Sudharsan, Josh Durkee