arXiv Machine Learning By Xu Lin (Tsinghua University, Beijing, China), Runheng Zuo (Tsinghua University, Beijing, China), Shengxuan Xu (Tsinghua University, Beijing, China), Qitai Tan (Tsinghua University, Beijing, China), Hongyu Lin (Tsinghua University, Beijing, China), Xiao-Ping Zhang (Tsinghua University, Beijing, China)

Beyond Average Error through Oracle-Informed Stress Tests for Time-Series Forecasting

Read the original on arXiv Machine Learning →

The paper introduces paired, mechanism‑controlled stress tests that decompose changes in expected squared error for time‑series forecasting into environmental risk and forecast‑oracle distance. Using an origin‑conditioned predictive oracle, the authors validate three end‑to‑end controls and apply the benchmark to 24 forecasters, revealing that many models exhibit higher realized MSE yet lower oracle distance under frequent switching, and that environmental risk dominates in most scenarios. The study also demonstrates that visually compelling discovery profiles often fail to replicate on independent data‑generating process realizations, underscoring the importance of component‑wise diagnosis and held‑out stability audits.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
2d ago

Beyond Model Ranking: Regime Diagnosis for Distributional-Statistical Misspecification in Industrial Time-Series Forecasting

The paper introduces a regime‑diagnosis framework for industrial time‑series forecasting, highlighting that canonical loss functions embed fixed statistical priors that are violated in real‑world demand regimes such as zero‑inflation, skewness, and high variability. It proposes the Regime‑wise Relative Bias Vector (RBV) as a metric‑agnostic diagnostic that decomposes bias into an intrinsic floor and an excess attributable to training. A large‑scale study across 13 loss objectives and 60,000+ series demonstrates that regime‑aware diagnosis distinguishes optimization‑from‑bias failures and that regime‑aware training can eliminate pooling‑induced bias that mere capacity scaling cannot.

By Pengyu Nie, Chenglang Xu, Yaoshi Chen, Chaogan Ren, Wei Hu, Chao Yang, Jiangong Zhang
arXiv Machine Learning
Sep 24

Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel

The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.

By Md Rezwanul Islam, Wael Mohammed