arXiv Machine Learning

Overlay\_dx - Automating forecasting evaluation

arXiv Machine Learning
Sep 14

Explaining Time Series Forecasting with Horizon-Resolved Attribution

The paper introduces Horizon-Resolved eXplanation (HRX), a framework that adds a horizon axis to time‑series forecasting explanations, allowing each forecast step to have its own importance map. HRX operates as a plug‑in for any differentiable forecaster, includes an evaluation protocol that tests the impact of removing top‑ranked inputs, and a rank criterion to decide when horizon resolution is beneficial. Experiments across multiple backbones and datasets demonstrate that incorporating the horizon axis improves explanation quality and that the step‑wise dependence is low‑dimensional, requiring only a few shared maps regardless of forecast length.

By Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, Wonbin Ahn
arXiv Machine Learning
Sep 4

Sliding-Window Reordering with Overlap Averaging: A Simple Time-Domain Augmentation for Multivariate Forecasting

The paper introduces a simple, model‑agnostic time‑domain augmentation called Sliding‑Window Reordering with Overlap Averaging. It transforms the joint input‑target sequence into overlapping windows, randomly reorders a fraction of them based on a variance criterion, and reconstructs the sequence by averaging overlaps to generate synthetic samples with controlled variation and minimal temporal distortion. Experiments show strong performance gains across nine long‑term forecasting benchmarks and four short‑term traffic benchmarks, with detailed ablations and diagnostics highlighting the effectiveness of each design choice.

By Jafar Bakhshaliyev, Johannes Burchert, Niels Landwehr, Lars Schmidt-Thieme
arXiv AI
Aug 17

Forecast Collapse in Time-Series Foundation Models

arXiv:2608. 14106v1 Announce Type: cross Abstract: When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation.

By Shu Wan, Miles Ma, Hank Zhu, Guangqi Liu, Stephen Wang, Qingsong Wen, Huan Liu
arXiv Machine Learning
Jun 30

fev-bench: A Realistic Benchmark for Time Series Forecasting

arXiv:2509. 26468v3 Announce Type: replace Abstract: Benchmark quality is critical for meaningful evaluation and sustained progress in time series forecasting, particularly with the rise of pretrained models.

By Oleksandr Shchur, Abdul Fatir Ansari, Caner Turkmen, Lorenzo Stella, Nick Erickson, Pablo Guerron, Michael Bohlke-Schneider, Yuyang Wang
arXiv AI
Aug 5

FinVerse: Financial Time-Series Benchmark

arXiv:2608. 03259v1 Announce Type: cross Abstract: As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become increasingly important.

By Jaehoon Lee, Jun Seo, Seunghan Lee, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Minjae Kim, Sungdong Yoo, Junhyeok Kang, Sangjun Han, Soonyoung Lee, Wonbin Ahn
arXiv AI
Aug 19

Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting

The paper critiques the prevalent use of mean squared error (MSE) for evaluating irregular time‑series forecasting, arguing that MSE is biased by timestamp sampling distributions. It introduces the Continuous‑time Squared Error (CSE), an importance‑weighted metric that theoretically offers a tighter asymptotic bound on continuous‑time risk than MSE. A comprehensive benchmark across synthetic, semi‑synthetic, and eight real‑world datasets demonstrates that CSE more accurately recovers continuous‑time risk, revealing limitations of relying solely on MSE.

By Rongwen Li, Haixin Xie, Xiao Wang, Changjian Chen