arXiv:2608. 10433v2 Announce Type: replace Abstract: Temporal reports are increasingly emitted alongside numerical forecasts and are often interpreted as statements about the computation producing those forecasts.
By Qipeng Qian, Yuntao Qian
arXiv:2606. 28670v1 Announce Type: cross Abstract: We introduce MACROCAST, a lightweight Time Series Foundation Model (TSFM) for real-time macroeconomic forecasting.
By Andrea Carriero, Davide Pettenuzzo, Shubhranshu Shekhar
arXiv:2608. 10433v4 Announce Type: replace Abstract: Time-series forecasters increasingly accompany numerical predictions with explicit temporal reports, such as delays or selected history, but a correct report need not describe the information actually used by the forecast.
By Qipeng Qian, Yuntao Qian
The paper introduces a regime‑diagnosis framework for industrial time‑series forecasting, highlighting that canonical loss functions embed fixed statistical priors that are violated in real‑world demand regimes such as zero‑inflation, skewness, and high variability. It proposes the Regime‑wise Relative Bias Vector (RBV) as a metric‑agnostic diagnostic that decomposes bias into an intrinsic floor and an excess attributable to training. A large‑scale study across 13 loss objectives and 60,000+ series demonstrates that regime‑aware diagnosis distinguishes optimization‑from‑bias failures and that regime‑aware training can eliminate pooling‑induced bias that mere capacity scaling cannot.
By Pengyu Nie, Chenglang Xu, Yaoshi Chen, Chaogan Ren, Wei Hu, Chao Yang, Jiangong Zhang
Forecast accuracy does not tell us which past inputs produced a prediction. We separate three questions for time-series models with known delay structure: can the true delay be recovered from the observed data, does the model report it, and does the forecast actually use the same history?
Time Series Foundation Models (TSFMs) have recently emerged as a highly promising paradigm for cross-domain zero-shot forecasting. However, existing evaluation protocols predominantly rely on static benchmarks with fixed historical test windows.
The paper argues that zero‑shot time‑series forecasting should be treated as an evidence‑access claim rather than merely a no‑parameter‑update condition. It introduces a source‑first taxonomy that distinguishes three evidence sources—frozen LLM prior reuse, parametric time‑series pretraining, and retrieval‑augmented external memory—from the architectures that implement them. The authors further outline four audit questions—task interface, forecast object and scoring, prediction‑time context, and resource budget—to make zero‑shot leaderboards transparent and comparable.
By Delun Kong, Wanyun Ling, Chenxi Liu, Ziyue Li
arXiv:2608. 10553v1 Announce Type: cross Abstract: Conformal prediction (CP) provides distribution-free prediction intervals for fixed forecasters, but its standard calibration procedure is often inefficient for time series data, where forecast errors are temporally dependent and change across time and operating conditions.
By Sangjin Jin, Kangmin Kim, Junhyeong Lee, Yongjae Lee
The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.
By Guangzhe Zhang
arXiv:2608. 10433v1 Announce Type: new Abstract: Forecast accuracy does not tell us which past inputs produced a prediction.
By Qipeng Qian, Yuntao Qian
LiveHouse-TS introduces an open‑world living benchmark for Time Series Foundation Models, evaluating them prequentially on real future data rather than static test windows. The benchmark captures continuous performance across seasonal changes, distribution shifts, and unexpected events, providing a more realistic assessment of model robustness. Experiments across 11 domains and 17 datasets show that model rankings can dramatically change under this live protocol.
By Haomin Wen, Ziyu Zhou, Qingxiang Liu, Siru Zhong, Yuxuan Liang
arXiv:2606. 27438v1 Announce Type: new Abstract: Since its initial release in 2020, Darts has become a widely used open-source Python library for time series analysis.
By Zhihao Dai, Dennis Bader, Alain Gysi