arXiv Machine Learning

Towards Reliable Zero-Shot Crowd Forecasting: Evaluating Time Series Foundation Models for Special Event Pedestrian Forecasting

arXiv:2607. 17758v1 Announce Type: new Abstract: Managing massive crowds during infrequent special events requires reliable real-time pedestrian-flow forecasting to ensure public safety and operational efficiency.

arXiv AI
Aug 19

LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models

LiveHouse-TS introduces an open‑world living benchmark for Time Series Foundation Models, evaluating them prequentially on real future data rather than static test windows. The benchmark captures continuous performance across seasonal changes, distribution shifts, and unexpected events, providing a more realistic assessment of model robustness. Experiments across 11 domains and 17 datasets show that model rankings can dramatically change under this live protocol.

By Haomin Wen, Ziyu Zhou, Qingxiang Liu, Siru Zhong, Yuxuan Liang
arXiv Machine Learning
Jul 28

Foundation Models and Fine-Tuning: Toward a New Generation of Models for Time Series Forecasting

arXiv:2607. 23146v1 Announce Type: new Abstract: Inspired by recent breakthroughs in large language models for natural language processing, foundation models have emerged as a promising paradigm for zero-shot time series forecasting, enabling accurate predictions on datasets never seen during pre-training.

By Morad Laglil, Bertrand Pracca, Emilie Devijver, Eric Gaussier
arXiv AI
Aug 24

RiskTraf: Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow Prediction

RiskTraf introduces a risk-extrapolated residual learning approach for multi-variate traffic flow prediction, leveraging raw flow, speed, and occupancy data from the new PEMSB-3V benchmark. The method freezes a trained spatio-temporal backbone and adds a lightweight residual head that learns from historical speed and occupancy to correct flow predictions across different traffic regimes. Experiments show consistent improvements over various backbones and outperform existing debiasing and distribution-shift adaptation techniques.

By Guangyu Wang, Zhidan Liu
arXiv AI
Aug 19

Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting

The paper critiques the prevalent use of mean squared error (MSE) for evaluating irregular time‑series forecasting, arguing that MSE is biased by timestamp sampling distributions. It introduces the Continuous‑time Squared Error (CSE), an importance‑weighted metric that theoretically offers a tighter asymptotic bound on continuous‑time risk than MSE. A comprehensive benchmark across synthetic, semi‑synthetic, and eight real‑world datasets demonstrates that CSE more accurately recovers continuous‑time risk, revealing limitations of relying solely on MSE.

By Rongwen Li, Haixin Xie, Xiao Wang, Changjian Chen