arXiv Machine Learning

fev-bench: A Realistic Benchmark for Time Series Forecasting

arXiv:2509. 26468v3 Announce Type: replace Abstract: Benchmark quality is critical for meaningful evaluation and sustained progress in time series forecasting, particularly with the rise of pretrained models.

arXiv Machine Learning
2d ago

The Nixtlaverse: An Open-Source Ecosystem for Forecasting

arXiv:2609.39741v1 Announce Type: new Abstract: Large forecasting applications often combine statistical, machine-learning, and neural models. These families solve the same problem but differ in fitt...

By Olivier Sprangers, Max Mergenthaler Canseco, Marco Peixeiro, Saul Caballero Ramirez, Mariana Menchero Garc\'ia, Jing-Qiang Goh, Han Wang, Nikhil Gupta, Rogelio Melo, Senbong Gee, Cristian Challu
arXiv AI
Aug 19

Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting

The paper critiques the prevalent use of mean squared error (MSE) for evaluating irregular time‑series forecasting, arguing that MSE is biased by timestamp sampling distributions. It introduces the Continuous‑time Squared Error (CSE), an importance‑weighted metric that theoretically offers a tighter asymptotic bound on continuous‑time risk than MSE. A comprehensive benchmark across synthetic, semi‑synthetic, and eight real‑world datasets demonstrates that CSE more accurately recovers continuous‑time risk, revealing limitations of relying solely on MSE.

By Rongwen Li, Haixin Xie, Xiao Wang, Changjian Chen
arXiv AI
Aug 19

LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models

LiveHouse-TS introduces an open‑world living benchmark for Time Series Foundation Models, evaluating them prequentially on real future data rather than static test windows. The benchmark captures continuous performance across seasonal changes, distribution shifts, and unexpected events, providing a more realistic assessment of model robustness. Experiments across 11 domains and 17 datasets show that model rankings can dramatically change under this live protocol.

By Haomin Wen, Ziyu Zhou, Qingxiang Liu, Siru Zhong, Yuxuan Liang