When Does Retrieval Help Time-Series Forecasting?
arXiv:2609. 20193v1 Announce Type: new Abstract: Retrieval plug-ins supply a deep forecaster with information its lookback window cannot carry.
arXiv:2607. 19383v1 Announce Type: cross Abstract: Pretrained generative foundation models cast forecasting as conditional generation from a learned predictive distribution and forecast unseen series zero-shot.
arXiv:2609. 20193v1 Announce Type: new Abstract: Retrieval plug-ins supply a deep forecaster with information its lookback window cannot carry.
arXiv:2609.24559v1 Announce Type: new Abstract: We present $t_0$, a family of open-weights foundation models for forecasting with multivariate context. We release its first two members: $\texttt{t0-a...
arXiv:2606. 09473v1 Announce Type: cross Abstract: Probabilistic forecasters are increasingly learned, yet the baselines they are compared against are often weak or omitted.
arXiv:2609.05561v1 Announce Type: cross Abstract: Rollcast is a probabilistic forecasting method for univariate time series that combines a compact set of rolling statistical anchors rather than rely...
arXiv:2609.06008v1 Announce Type: cross Abstract: We present Cadence, an error-bounded lossy compressor for numeric time series pairing a 330M-parameter time-series foundation model (Google TimesFM-3...
arXiv:2607. 12248v1 Announce Type: cross Abstract: Large pretrained time-series models such as TimesFM are attractive for financial forecasting, but raw directional accuracy is a misleading scoreboard in equity markets.
arXiv:2608. 14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief.
arXiv:2609.39386v1 Announce Type: new Abstract: Pretrained time-series foundation models (TSFMs) are evaluated as forecasters of future values, yet for sparse series many decisions depend only on whi...
The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.
The paper introduces paired, mechanism‑controlled stress tests that decompose changes in expected squared error for time‑series forecasting into environmental risk and forecast‑oracle distance. Using an origin‑conditioned predictive oracle, the authors validate three end‑to‑end controls and apply the benchmark to 24 forecasters, revealing that many models exhibit higher realized MSE yet lower oracle distance under frequent switching, and that environmental risk dominates in most scenarios. The study also demonstrates that visually compelling discovery profiles often fail to replicate on independent data‑generating process realizations, underscoring the importance of component‑wise diagnosis and held‑out stability audits.
Time-series foundation models are evaluated almost exclusively on public archives that predate them, so a strong score cannot be separated from having seen the test set during pretraining. The obvious...
arXiv:2607. 07951v1 Announce Type: new Abstract: Wildfire smoke events produce extreme PM$_{2.