arXiv Machine Learning

From Seasonality to Semantics: Benchmarking a Hybrid Probabilistic Forecasting System for Roadblocks in Bolivia

arXiv:2607. 21785v1 Announce Type: cross Abstract: Roadblocks in Bolivia are a social conflict phenomenon with devastating economic impacts, estimated at losses equivalent to 4% of the national Gross Domestic Product.

arXiv AI
Aug 25

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

arXiv:2608.23058v1 Announce Type: new Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external too...

By Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu, Wei Wei, Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng
arXiv AI
Sep 12

When Does Text Inform? Benchmarking Information-Theoretic Metrics for Multimodal Time-Series Forecasting

The paper introduces a synthetic benchmark for multimodal time‑series forecasting that evaluates how well text annotations contribute to predictions. By generating controlled signals with semantically correct, incorrect, and irrelevant annotations, the authors can precisely measure the true information content. Six mutual‑information estimators (KSG, MINE, InfoNCE, CCA, PID, and V‑information) are tested, all correctly ranking useful annotations and enabling annotation auditing without model training. The benchmark also highlights each estimator’s limitations and validates findings on seven real datasets, providing practical guidelines for metric implementation.

By Emma Andrews, Gianmarco Mengaldo
arXiv Machine Learning
Sep 7

Evaluating Large Language Models for Forced Outage Risk Prediction: Benefits and Comparison to Machine Learning

This study evaluates large language models (LLMs) for predicting weather‑related forced outage risk in a distribution grid using a zero‑shot approach without labeled training data. The task is framed as binary severity classification over 3h, 6h, and 12h horizons, leveraging six years of outage records and high‑resolution weather data from central Texas. Four zero‑shot LLMs are compared to two supervised classifiers under two input settings—current weather observations and forecast data—showing that supervised models lead on macro‑F1 and precision, while newer LLMs achieve competitive scores and offer complementary strengths in reasoning and geographic scalability.

By Christos Petridis, Zoran Obradovic, Mladen Kezunovic
arXiv Machine Learning
Aug 19

Proactive Road Safety Intervention in Australia: Predicting Risky Driving Hotspots from Connected Vehicle Data

The paper proposes a proactive approach to road safety in Greater Sydney by using connected vehicle telemetry to predict risky driving events before crashes occur. It quantifies risky driving with g‑force thresholds and builds spatio‑temporal heatmaps to locate high‑risk zones. Eight predictive models were compared, with ARIMA achieving the lowest error and showing that simple time‑series methods can rival deep learning when data are limited, highlighting the value of IoT data for targeted safety interventions.

By Adriana-Simona Mih\u{a}i\c{t}\u{a}, Clarence Cheung, Artur Grigorev, Tuo Mao, David Lillo-Trynes
arXiv Machine Learning
Aug 27

Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings

The Planetary Prediction Engine (PPE) is an autonomous AI system that transforms natural-language queries into end-to-end geospatial predictions. It automatically retrieves and fuses multimodal datasets from open-web and Earth observation sources, incorporates foundation model embeddings, and searches task‑specific model families with overfitting safeguards. Across multiple domains, PPE outperforms state‑of‑the‑art baselines, improving regression metrics for CDC health indicators, FEMA risk indices, and the Social Vulnerability Index, doubling accuracy for Nigerian food security indicators, and achieving higher recall in Ebola outbreak nowcasting.

By Evelyn Ma, Rama Kumar Pasumarthi, Kishwar Shafin, Mandar Sharma, Mimi Sun, Hamed Sadeghi, Dav M. Ebengo, Mbulayi Onesime, Rouslan Solomakhin, John Wamburu, William Ogallo, Aisha Walcott-Bryant, Sanxing Chen, Arbaaz Muslim, Yael Mayer, Ronald Ho, Roy Lee, Ruth Alcantara, Abdoulaye Diack, Monica Bharel, Lambert Rosique, Jeremy Amez-Droz, Christopher Haire, James Manyika, Yossi Matias, Niv Efron, Gautam Prasad, Shravya Shetty
arXiv AI
Aug 18

ReasonCast: Agentic Demand Forecasting with Selective Semantic Reasoning

arXiv:2608. 15291v1 Announce Type: new Abstract: Demand forecasting increasingly requires combining two complementary sources of information: historical sales reveal recurring numerical dynamics, while future promotions, holidays, price changes, and platform interventions provide forward-looking knowledge.

By Ziyue Yang, Chaolin Xu, Yijing Wang, Tiankai Gu, Hui Yang, Yanhong Lin, Kaiyuan Liu, Fei Xiao
arXiv Machine Learning
Sep 25

Improving global precipitation forecasts with an AI weather model trained on satellite observations

The paper presents Laxmi, a retrained version of the AIFS weather model that uses satellite-based precipitation observations instead of ERA5 reanalysis data. Laxmi achieves a 19% improvement in global probabilistic accuracy, reduces drizzle overprediction by 33%, and boosts the 95th percentile Brier skill score by 57%. In a case study of 10 Indian tropical storms, Laxmi accurately forecasted 150 mm event-total precipitation in 7 events, outperforming both the original AIFS and the leading physical model IFS.

By Julian F. Schmitt, Bertrand Delorme, Robert C. King, Yashica Patodia, Tapio Schneider, Aditi Sheshadri, Ravi Jain