arXiv AI

Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools

The paper introduces FWBench, a benchmark for assessing how language models choose and employ time‑series forecasts under budget constraints. Using 1,251 electricity and cycle‑hire cases, the study evaluates agents that select models, histories, and horizons to submit capacities that minimize a loss‑cost objective. Experiments with both hosted and local configurations—including small language models and TSFMs—show that GPT‑6 Astra can achieve superior performance by selectively buying inexpensive short‑horizon forecasts, using only 2.5% of the budget.

arXiv AI
Aug 25

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

arXiv:2608.23058v1 Announce Type: new Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external too...

By Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu, Wei Wei, Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng