The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.
By Md Rezwanul Islam, Wael Mohammed
arXiv:2607. 17511v1 Announce Type: new Abstract: Large \emph{Time Series Foundation Models} (TSFMs) demonstrate strong zero-shot forecasting capabilities across diverse domains.
By Wentao Gao, Jiuyong Li, Lin Liu, Thuc Duy Le, Jixue Liu, Yanchang Zhao, Yun Chen
arXiv:2607. 05450v1 Announce Type: cross Abstract: This paper explores the "Granularity Paradox" in time-series forecasting, wherein finer temporal disaggregation (e.
By Hugo Moreira
arXiv:2607. 19659v1 Announce Type: new Abstract: Time-series foundation models can forecast across heterogeneous domains without task-specific training, but their forecasts are fixed once produced and cannot directly incorporate task-specific expert feedback.
By Hung Le, Minh Hoang Nguyen, Manh Nguyen, Huu Hiep Nguyen, Dai Do
The paper introduces Generalized Gibbs Ensemble Weighting (GGEW), a probabilistic framework that assigns weights to forecasting models using a Gibbs-style exponential transformation of normalized predictive loss. GGEW extends basic weighting through numerical stabilization, diversity-aware score corrections, and online hyperparameter adaptation, yielding variants such as Stable Gibbs weighting, Directional Gibbs-NCL, and Symmetric Gibbs-NCL. The authors evaluate GGEW on M4 competition submissions and real-world datasets (Monash Traffic, Electricity, Solar), finding that Gibbs-style adaptive weighting is competitive across various settings, though performance varies by dataset, horizon, and deployment protocol.
By Prasen R. Nuthanakaluva, Nava K. Gaddam
arXiv:2508. 09904v3 Announce Type: replace-cross Abstract: Real-world forecasting requires models to integrate not only historical data but also relevant contextual information provided in textual form.
By Arjun Ashok, Andrew Robert Williams, Vincent Zhihao Zheng, Irina Rish, Nicolas Chapados, \'Etienne Marcotte, Valentina Zantedeschi, Alexandre Drouin