The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.
By Md Rezwanul Islam, Wael Mohammed
The paper introduces loss‑conditioned state execution, a model‑agnostic technique that decides whether to apply a world model’s proposed state change or keep the current state based on whether the change reduces downstream loss. It formalizes state movability as the existence of a loss‑reducing feasible correction and constructs loss‑specific proposals from predictive distributions, executing them only when a groupwise lower confidence bound on loss improvement is positive. Experiments on forecasting and dynamics benchmarks show that the method accepts updates for a subset of cases, achieving lower bounded loss than persistence or always executing the proposal, and highlights that event predictability and loss‑based decisions must be evaluated separately.
By Jintao Xu, Zhengyu Chen, Ben Zhang, Yongzhi Qi, Jianshen Zhang
arXiv:2606. 24996v1 Announce Type: new Abstract: Forecasting leaderboards rank models by predictive quality, but their winners are often read as deployment-ready top-1 advice.
By Geumyoung Kim
The paper introduces a counterfactual tool ranking framework that accounts for authority, historical support, and estimation nuances. Using eleven enterprise-inspired tools, synthetic and real-world experiments on the Berkeley Function Calling Leaderboard, the study compares direct regression and doubly robust (DR) methods, finding that DR performs better in shifted environments while direct regression excels in linear settings. The authors also evaluate Qwen2.5 models on held-out tasks, analyze policy differences under missing support, and present a falsifiable evaluation method with publicly available evidence.
By Jiapeng Li
The paper introduces a regime‑diagnosis framework for industrial time‑series forecasting, highlighting that canonical loss functions embed fixed statistical priors that are violated in real‑world demand regimes such as zero‑inflation, skewness, and high variability. It proposes the Regime‑wise Relative Bias Vector (RBV) as a metric‑agnostic diagnostic that decomposes bias into an intrinsic floor and an excess attributable to training. A large‑scale study across 13 loss objectives and 60,000+ series demonstrates that regime‑aware diagnosis distinguishes optimization‑from‑bias failures and that regime‑aware training can eliminate pooling‑induced bias that mere capacity scaling cannot.
By Pengyu Nie, Chenglang Xu, Yaoshi Chen, Chaogan Ren, Wei Hu, Chao Yang, Jiangong Zhang
arXiv:2606. 04342v1 Announce Type: cross Abstract: Multi-step time series forecasting (MSF) is commonly evaluated using point-wise error metrics such as mean squared error (MSE), implicitly treating the conditional mean as a sufficient target.
By Riku Green, Zahraa S. Abdallah, Telmo M Silva Filho