arXiv:2609.36966v1 Announce Type: cross
Abstract: Covariate effects vary across contexts and shift over time, requiring forecasters to assess how to use them for each forecasting context. As forecast...
By Donguk Kwon, Wooseok Jeong, Dongha Lee
arXiv:2608. 14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief.
By Fabricio F Costa
The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.
By Md Rezwanul Islam, Wael Mohammed
RATL is a plug‑in method for multivariate time‑series forecasting that uses a frozen base forecaster to build a memory of its historical forecast residuals. During inference, RATL retrieves residual trajectories from similar past contexts and employs a set‑aware router to combine them, providing learned feedback correction. Experiments demonstrate that this residual‑retrieval approach improves the performance of the base forecaster across various benchmarks and backbones.
By Yuchen He, Yueyang Cang, Zhiyuan Ning, Ningyu Wang, Li Shi
The paper demonstrates that open‑ended Theory‑of‑Mind trackers can produce valid beliefs that are absent from finite reference sets, and that treating unmatched outputs as false can reverse model‑selection rankings. By recoding references for 259 beliefs, the authors show a dramatic drop in weighted prevalence and a reversal of strictly proper Brier risk, with similar distortions observed in a 301‑question NQ‑open DPR‑BERT pipeline. The study further reveals that 90‑96% of audited unmatched beliefs are literally true, and introduces a TriSource‑Restore method that anchors reference labels to a probability‑sampled human pilot to restore calibration and ranking integrity.
By Zhexi Feng, Wuxi Chen, Bingrui Zhang
arXiv:2609.39386v1 Announce Type: new
Abstract: Pretrained time-series foundation models (TSFMs) are evaluated as forecasters of future values, yet for sparse series many decisions depend only on whi...
By Daniel Schoess, Florian von Wangenheim
The paper introduces Horizon-Resolved eXplanation (HRX), a framework that adds a horizon axis to time‑series forecasting explanations, allowing each forecast step to have its own importance map. HRX operates as a plug‑in for any differentiable forecaster, includes an evaluation protocol that tests the impact of removing top‑ranked inputs, and a rank criterion to decide when horizon resolution is beneficial. Experiments across multiple backbones and datasets demonstrate that incorporating the horizon axis improves explanation quality and that the step‑wise dependence is low‑dimensional, requiring only a few shared maps regardless of forecast length.
By Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, Wonbin Ahn
Forecast-Dojo is a replayable environment designed to benchmark and train large language model (LLM) forecasting agents. It integrates resolved prediction‑market questions with dated news, enabling agents to research events and revisit predictions at successive historical dates. The platform includes 1,568 Polymarket events, 18.8 million dated news articles, and supports repeated evaluation, training interactions, and outcome feedback, with evidence that research tools lower Brier scores across 12 tested models, though all models still lag behind historical market forecasts.
By Liqin Ye, Haorui Wang, Fardin Ahmed, Rongzhi Zhang, Yuan He, Ziyuan Lin, Yanbin Yin, Jing Peng, Michael Galarnyk, Sudheer Chava, Chao Zhang
arXiv:2609.35917v1 Announce Type: new
Abstract: Graph neural networks are increasingly applied to road-level crash prediction, but the stability of their reported gains has received little scrutiny....
By Maurya Patel
arXiv:2603.04275v2 Announce Type: replace-cross
Abstract: We introduce inference methods for score decompositions, which partition scoring functions for predictive assessment into three interpretable...
By Timo Dimitriadis, Marius Puke
arXiv:2608. 10433v4 Announce Type: replace Abstract: Time-series forecasters increasingly accompany numerical predictions with explicit temporal reports, such as delays or selected history, but a correct report need not describe the information actually used by the forecast.
By Qipeng Qian, Yuntao Qian
arXiv:2608. 06765v1 Announce Type: new Abstract: Continuous-time dynamic graph models predict future links by compressing past interactions into neural states.
By Minwoo Yu, Young-guk Ha