Towards Data Science

Five Ways to Fine-Tune Chronos-2, the Time Series Foundation Model

In Part 1 of this series, we introduced Chronos-2, a time-series foundation model. We got our hands dirty by walking through a real case study and saw what Chronos-2 can do straight out of the box, with no training.

arXiv Machine Learning
Jun 5

Toto 2.0: Time Series Forecasting Enters the Scaling Era

arXiv:2605. 20119v2 Announce Type: replace Abstract: We show that time series foundation models scale: a single training recipe produces reliable forecast-quality improvements from 4M to 2.

By Emaad Khwaja, Chris Lettieri, Gerald Woo, Eden Belouadah, Marc Cenac, Guillaume Jarry, Enguerrand Paquin, Xunyi Zhao, Viktoriya Zhukov, Othmane Abou-Amal, Chenghao Liu, Ameet Talwalkar, David Asker
arXiv Machine Learning
1d ago

Foundations without Fundamentals: Zero-Shot Blind Spots in Time Series FMs

The paper introduces SimpleTimeBench, a diagnostic suite testing basic temporal primitives like monotonic trends, periodic signals, and leading indicator covariates. It finds that prominent multivariate Time Series Foundation Models (Chronos‑2, Moirai, and Toto) often produce suboptimal zero‑shot forecasts for these simple patterns, and that fine‑tuning can improve specific tasks while harming performance on other fundamentals. These failures persist in real‑world sensor forecasting, indicating that current TSFMs may lack the inductive biases needed to capture straightforward relationships, thereby limiting their practical reliability.

By Nafiseh Ghoroghchian, Haipeng Zhang, Shuyi Han, Alex Labach, George Stein
arXiv Machine Learning
Sep 25

Time-Series Foundation Models That Understand Data Revisions

The paper introduces VINTAGE-TS, a revision‑aware time‑series foundation model that separates observation time from information‑availability time. It predicts both the next period’s first‑published value and the value available after a fixed delay, maintaining a joint distribution to capture their dependence and uncertainty. The authors provide a detailed evaluation protocol, software tools for validity‑interval reconstruction and delayed‑label filtering, and a synthetic demonstration with a 25‑configuration sensitivity suite to illustrate performance variability and the impact of hindsight contamination.

By Taimoor Ahmad
arXiv Machine Learning
Jun 9

Zero and Few Shot Load Forecasting with Large Language Models

arXiv:2411. 11350v2 Announce Type: replace Abstract: Deep learning models have shown strong performance in load forecasting, but they generally require large amounts of data for model training before being applied to new scenarios, which limits their effectiveness in data-scarce scenarios.

By Wenlong Liao, Chengrui Zhang, Zhe Yang, Mengshuo Jia, Christian Rehtanz, Jiannong Fang, Fernando Port\'e-Agel
arXiv Machine Learning
Sep 21

Tracing the Evidence Behind Zero-Shot Time-Series Forecasting: A Source-First Taxonomy and Audit Framework

The paper argues that zero‑shot time‑series forecasting should be treated as an evidence‑access claim rather than merely a no‑parameter‑update condition. It introduces a source‑first taxonomy that distinguishes three evidence sources—frozen LLM prior reuse, parametric time‑series pretraining, and retrieval‑augmented external memory—from the architectures that implement them. The authors further outline four audit questions—task interface, forecast object and scoring, prediction‑time context, and resource budget—to make zero‑shot leaderboards transparent and comparable.

By Delun Kong, Wanyun Ling, Chenxi Liu, Ziyue Li