arXiv Machine Learning

Beyond Average Error through Oracle-Informed Stress Tests for Time-Series Forecasting

The paper introduces paired, mechanism‑controlled stress tests that decompose changes in expected squared error for time‑series forecasting into environmental risk and forecast‑oracle distance. Using an origin‑conditioned predictive oracle, the authors validate three end‑to‑end controls and apply the benchmark to 24 forecasters, revealing that many models exhibit higher realized MSE yet lower oracle distance under frequent switching, and that environmental risk dominates in most scenarios. The study also demonstrates that visually compelling discovery profiles often fail to replicate on independent data‑generating process realizations, underscoring the importance of component‑wise diagnosis and held‑out stability audits.

arXiv Machine Learning
2d ago

Beyond Model Ranking: Regime Diagnosis for Distributional-Statistical Misspecification in Industrial Time-Series Forecasting

The paper introduces a regime‑diagnosis framework for industrial time‑series forecasting, highlighting that canonical loss functions embed fixed statistical priors that are violated in real‑world demand regimes such as zero‑inflation, skewness, and high variability. It proposes the Regime‑wise Relative Bias Vector (RBV) as a metric‑agnostic diagnostic that decomposes bias into an intrinsic floor and an excess attributable to training. A large‑scale study across 13 loss objectives and 60,000+ series demonstrates that regime‑aware diagnosis distinguishes optimization‑from‑bias failures and that regime‑aware training can eliminate pooling‑induced bias that mere capacity scaling cannot.

By Pengyu Nie, Chenglang Xu, Yaoshi Chen, Chaogan Ren, Wei Hu, Chao Yang, Jiangong Zhang
arXiv Machine Learning
Sep 24

Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel

The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.

By Md Rezwanul Islam, Wael Mohammed
arXiv Machine Learning
Sep 25

Time-Series Foundation Models That Understand Data Revisions

The paper introduces VINTAGE-TS, a revision‑aware time‑series foundation model that separates observation time from information‑availability time. It predicts both the next period’s first‑published value and the value available after a fixed delay, maintaining a joint distribution to capture their dependence and uncertainty. The authors provide a detailed evaluation protocol, software tools for validity‑interval reconstruction and delayed‑label filtering, and a synthetic demonstration with a 25‑configuration sensitivity suite to illustrate performance variability and the impact of hindsight contamination.

By Taimoor Ahmad
arXiv Machine Learning
Aug 27

When Does Context Routing Help? A Systematic Study of Multi-Modal Fusion in Time Series Forecasting

The paper investigates when auxiliary context can genuinely improve multi‑modal time series forecasting. It identifies two necessary dataset‑level conditions: the target must not be dominated by a last‑value shortcut (low autocorrelation) and the context must provide additional information beyond history (non‑zero conditional mutual information). Experiments on a large mixture‑of‑experts model and several fusion mechanisms show that only when both conditions hold does context routing yield a substantial reduction in mean‑squared error; otherwise its contribution collapses to a capacity floor.

By Ruizhe Zhou, Gaoyuan Du, Xiaoyang Liu, Haoqi Yao, Deepayan Chakrabarti, Jiating Lin, Yixuan Shen