arXiv AI

LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models

LiveHouse-TS introduces an open‑world living benchmark for Time Series Foundation Models, evaluating them prequentially on real future data rather than static test windows. The benchmark captures continuous performance across seasonal changes, distribution shifts, and unexpected events, providing a more realistic assessment of model robustness. Experiments across 11 domains and 17 datasets show that model rankings can dramatically change under this live protocol.

arXiv Machine Learning
Jun 30

fev-bench: A Realistic Benchmark for Time Series Forecasting

arXiv:2509. 26468v3 Announce Type: replace Abstract: Benchmark quality is critical for meaningful evaluation and sustained progress in time series forecasting, particularly with the rise of pretrained models.

By Oleksandr Shchur, Abdul Fatir Ansari, Caner Turkmen, Lorenzo Stella, Nick Erickson, Pablo Guerron, Michael Bohlke-Schneider, Yuyang Wang
arXiv Machine Learning
Jun 5

Toto 2.0: Time Series Forecasting Enters the Scaling Era

arXiv:2605. 20119v2 Announce Type: replace Abstract: We show that time series foundation models scale: a single training recipe produces reliable forecast-quality improvements from 4M to 2.

By Emaad Khwaja, Chris Lettieri, Gerald Woo, Eden Belouadah, Marc Cenac, Guillaume Jarry, Enguerrand Paquin, Xunyi Zhao, Viktoriya Zhukov, Othmane Abou-Amal, Chenghao Liu, Ameet Talwalkar, David Asker
arXiv Machine Learning
Jun 18

Benchmarking Physics-Informed Time-Series Models for Operational Global Station Weather Forecasting

arXiv:2406. 14399v4 Announce Type: replace Abstract: The development of Time-Series Forecasting (TSF) models is often constrained by the lack of comprehensive datasets, especially in Global Station Weather Forecasting (GSWF), where existing datasets are small, temporally short, and spatially sparse.

By Tao Han, Zhibin Wen, Zhenghao Chen, Dazhao Du, Song Guo, Lei Bai