Forecast accuracy does not tell us which past inputs produced a prediction. We separate three questions for time-series models with known delay structure: can the true delay be recovered from the observed data, does the model report it, and does the forecast actually use the same history?
arXiv:2608. 10433v4 Announce Type: replace Abstract: Time-series forecasters increasingly accompany numerical predictions with explicit temporal reports, such as delays or selected history, but a correct report need not describe the information actually used by the forecast.
By Qipeng Qian, Yuntao Qian
arXiv:2608. 10433v2 Announce Type: replace Abstract: Temporal reports are increasingly emitted alongside numerical forecasts and are often interpreted as statements about the computation producing those forecasts.
By Qipeng Qian, Yuntao Qian
arXiv:2606. 28670v1 Announce Type: cross Abstract: We introduce MACROCAST, a lightweight Time Series Foundation Model (TSFM) for real-time macroeconomic forecasting.
By Andrea Carriero, Davide Pettenuzzo, Shubhranshu Shekhar
arXiv:2607. 10972v1 Announce Type: new Abstract: Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop.
By Aleh Manchuliantsau
The paper introduces paired, mechanism‑controlled stress tests that decompose changes in expected squared error for time‑series forecasting into environmental risk and forecast‑oracle distance. Using an origin‑conditioned predictive oracle, the authors validate three end‑to‑end controls and apply the benchmark to 24 forecasters, revealing that many models exhibit higher realized MSE yet lower oracle distance under frequent switching, and that environmental risk dominates in most scenarios. The study also demonstrates that visually compelling discovery profiles often fail to replicate on independent data‑generating process realizations, underscoring the importance of component‑wise diagnosis and held‑out stability audits.
By Xu Lin (Tsinghua University, Beijing, China), Runheng Zuo (Tsinghua University, Beijing, China), Shengxuan Xu (Tsinghua University, Beijing, China), Qitai Tan (Tsinghua University, Beijing, China), Hongyu Lin (Tsinghua University, Beijing, China), Xiao-Ping Zhang (Tsinghua University, Beijing, China)
The paper introduces VINTAGE-TS, a revision‑aware time‑series foundation model that separates observation time from information‑availability time. It predicts both the next period’s first‑published value and the value available after a fixed delay, maintaining a joint distribution to capture their dependence and uncertainty. The authors provide a detailed evaluation protocol, software tools for validity‑interval reconstruction and delayed‑label filtering, and a synthetic demonstration with a 25‑configuration sensitivity suite to illustrate performance variability and the impact of hindsight contamination.
By Taimoor Ahmad
The paper introduces a formal framework and benchmark for time‑series world models (TSWMs) that separates state, actions, and exogenous inputs, and defines a new metric called mechanism consistency to evaluate whether model predictions move in the expected direction when actions change. Experiments on eight public datasets show that using a frozen latent prediction space and gated output fusion improves prediction accuracy, while prediction error and mechanism consistency often diverge, with the best‑performing models sometimes failing to exhibit consistent directional responses. Adding a directional supervision loss significantly boosts mechanism consistency without affecting mean‑absolute error, providing a practical recipe for building more reliable TSWMs.
By Haochen Zhang, Jiaheng Guo, Zhen Xu, Zachary Plotkin, Nicholas Konz, Zhen Tan, Tianlong Chen
The paper proposes a method for controlling downside risk when adjusting forecasts from frozen models, such as foundation models, by combining a static corrector and an online corrector on the simplex. Using only post‑horizon losses, the approach achieves minimal deterioration (0.15%) and up to 11.5% gains across 28 forecast pairs, and consistently reduces mean MSE in day‑ahead load forecasts for seven European bidding zones. The method’s applicability is bounded by three empirical conditions related to expert speed, stream length, and outcome alignment.
By Minkyoung Kim, Hyunjung Byun, Yohan Lee, Beakcheol Jang
RATL is a plug‑in method for multivariate time‑series forecasting that uses a frozen base forecaster to build a memory of its historical forecast residuals. During inference, RATL retrieves residual trajectories from similar past contexts and employs a set‑aware router to combine them, providing learned feedback correction. Experiments demonstrate that this residual‑retrieval approach improves the performance of the base forecaster across various benchmarks and backbones.
By Yuchen He, Yueyang Cang, Zhiyuan Ning, Ningyu Wang, Li Shi
Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop. We study the complementary setting in which ground truth is delayed, censored, or private, so deterministic code cannot check correctness at scoring time and must instead issue a code-owned provisional forecast.
The paper investigates when forecasting accuracy can reliably reveal the underlying temporal structure of a time series. It shows that a small forecast margin does not automatically mean structural ambiguity and introduces a stability-based measure that assesses how well different temporal mechanisms can be distinguished given uncertainty in the selection objective. Experiments demonstrate that this stability metric better predicts when forecast-only structural selection succeeds or fails compared to relying solely on forecast margin.
By Qipeng Qian, Yuntao Qian