arXiv Machine Learning

Trend strength predicts when generative foundation models win: a power-controlled benchmark, a mechanism, and an actionable selection rule

arXiv:2607. 19383v1 Announce Type: cross Abstract: Pretrained generative foundation models cast forecasting as conditional generation from a learned predictive distribution and forecast unseen series zero-shot.

arXiv Machine Learning
Sep 24

Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel

The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.

By Md Rezwanul Islam, Wael Mohammed
arXiv Machine Learning
Sep 22

Beyond Average Error through Oracle-Informed Stress Tests for Time-Series Forecasting

The paper introduces paired, mechanism‑controlled stress tests that decompose changes in expected squared error for time‑series forecasting into environmental risk and forecast‑oracle distance. Using an origin‑conditioned predictive oracle, the authors validate three end‑to‑end controls and apply the benchmark to 24 forecasters, revealing that many models exhibit higher realized MSE yet lower oracle distance under frequent switching, and that environmental risk dominates in most scenarios. The study also demonstrates that visually compelling discovery profiles often fail to replicate on independent data‑generating process realizations, underscoring the importance of component‑wise diagnosis and held‑out stability audits.

By Xu Lin (Tsinghua University, Beijing, China), Runheng Zuo (Tsinghua University, Beijing, China), Shengxuan Xu (Tsinghua University, Beijing, China), Qitai Tan (Tsinghua University, Beijing, China), Hongyu Lin (Tsinghua University, Beijing, China), Xiao-Ping Zhang (Tsinghua University, Beijing, China)