When Does Retrieval Help Time-Series Forecasting?
arXiv:2609. 20193v1 Announce Type: new Abstract: Retrieval plug-ins supply a deep forecaster with information its lookback window cannot carry.
arXiv:2607. 12248v1 Announce Type: cross Abstract: Large pretrained time-series models such as TimesFM are attractive for financial forecasting, but raw directional accuracy is a misleading scoreboard in equity markets.
arXiv:2609. 20193v1 Announce Type: new Abstract: Retrieval plug-ins supply a deep forecaster with information its lookback window cannot carry.
The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.
arXiv:2502. 18834v3 Announce Type: replace-cross Abstract: Financial time series (FinTS) record the behavior of human-brain-augmented decision-making, capturing valuable historical information that can be leveraged for profitable investment strategies.
arXiv:2608. 03259v1 Announce Type: cross Abstract: As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become increasingly important.
UQ-LOB is a lightweight, encoder‑agnostic module that adds uncertainty quantification to any pretrained limit order book (LOB) encoder. It offers two variants: UQ‑regression, which outputs a calibrated Gaussian over future tick displacement, and UQ‑classification, which outputs a categorical distribution over down/up/stationary. On 5.2 billion LOB events across seven cryptocurrency assets, UQ‑regression achieves near‑nominal 68 % interval coverage, and selecting the top 10 % most confident predictions boosts directional macro F1 by 0.11–0.15 for regression and 0.05–0.11 for classification, reaching F1 scores of 0.88 (down) and 0.83 (up) at a 5‑second horizon.
arXiv:2606. 28670v1 Announce Type: cross Abstract: We introduce MACROCAST, a lightweight Time Series Foundation Model (TSFM) for real-time macroeconomic forecasting.
Time-series foundation models are evaluated almost exclusively on public archives that predate them, so a strong score cannot be separated from having seen the test set during pretraining. The obvious...
arXiv:2608. 14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief.
arXiv:2608. 08825v1 Announce Type: cross Abstract: Foundation models for time series forecasting demonstrate impressive zero-shot generalization but often underperform on specialized domains such as high-frequency finance.
arXiv:2609.39386v1 Announce Type: new Abstract: Pretrained time-series foundation models (TSFMs) are evaluated as forecasters of future values, yet for sparse series many decisions depend only on whi...
arXiv:2608. 14106v1 Announce Type: cross Abstract: When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation.
arXiv:2605.24564v2 Announce Type: replace Abstract: Backtesting large language models (LLMs) on historical financial data is unreliable when their pre-training data include the evaluated events. An L...