arXiv AI

TS-Memory: Plug-and-Play Memory for Time Series Foundation Models

arXiv:2602. 11550v2 Announce Type: replace-cross Abstract: Time Series Foundation Models (TSFMs) achieve strong zero-shot forecasting through large-scale pre-training, but adapting them to downstream domains under distribution shift remains challenging.

arXiv Machine Learning
5d ago

Aurora-X: Built for Extreme Time Series Forecasting

Aurora‑X is a billion‑parameter time‑series foundation model designed for extreme forecasting tasks. It employs a progressive curriculum that starts with channel‑independent pretraining, then adds cross‑variable dependencies, variable context and horizon lengths, and optional future covariates during mid‑training. A variable‑resolution post‑training stage allows adjustable temporal spans per token at inference, while a pattern‑guided mixture‑of‑experts expands capacity through sparse activation and expert specialization. An implicit quantile network head predicts arbitrary quantiles, enhancing probabilistic forecasting flexibility. Experiments on GIFT‑Eval, TIME, FEV‑Bench, TFB, and DAG‑Bench show state‑of‑the‑art performance against both pretrained TSFMs and task‑specific supervised models.

By Xingjian Wu, Chenjuan Guo, Xiangfei Qiu, Zhigang Hu, Hanyin Cheng, Peng Chen, Yang Shu, Jilin Hu, Bin Yang
arXiv Machine Learning
Sep 15

Tabby: An Open Pretraining Recipe for Time Series Foundation Models

arXiv:2609.13956v1 Announce Type: new Abstract: In this report, we release Tabby, a long context probabilistic time series foundation model, together with a complete and open recipe of how it was bui...

By Shifeng Xie, Bahaeddine Abdessalem, Zehao Xiao, Youssef Attia El Hili, Ambroise Odonnat, Zhiwei Dong, Lei Zan, Themis Palpanas, Jianfeng Zhang, Lujia Pan, Keli Zhang, Malik Tiomoko
arXiv AI
Sep 4

RATL: Learning from Retrieved Residuals for Robust Multivariate Time-Series Forecasting

RATL is a plug‑in method for multivariate time‑series forecasting that uses a frozen base forecaster to build a memory of its historical forecast residuals. During inference, RATL retrieves residual trajectories from similar past contexts and employs a set‑aware router to combine them, providing learned feedback correction. Experiments demonstrate that this residual‑retrieval approach improves the performance of the base forecaster across various benchmarks and backbones.

By Yuchen He, Yueyang Cang, Zhiyuan Ning, Ningyu Wang, Li Shi
arXiv Machine Learning
Aug 5

Maglev: Sliding Recurrent Memory

arXiv:2608. 02870v1 Announce Type: new Abstract: We introduce \ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training.

By Bo Liu, Qiang Liu
arXiv Machine Learning
Aug 19

Continuous Evolution Pool: Taming Recurring Concept Drift in Online Time Series Forecasting

The paper introduces the Continuous Evolution Pool (CEP), a replay‑free framework for online time series forecasting that tackles recurring concept drift. CEP maintains a dynamic pool of specialized forecasters, using lightweight statistical genes to identify concepts, spawn new models when distribution shifts occur, and prune obsolete ones under memory limits. Experiments on real‑world datasets show CEP reduces forecasting error by up to 24% compared to state‑of‑the‑art baselines, especially in scenarios with pronounced recurring drift.

By Tianxiang Zhan, Ming Jin, Yuanpeng He, Yuxuan Liang, Shirui Pan
arXiv Machine Learning
Sep 22

A Hybrid Attention Model Learning Unified Time-aware Patch Representation for Irregular Multivariate Time Series Forecasting

The paper introduces a hybrid attention model that learns a unified time‑aware patch representation for irregular multivariate time series (IMTS) forecasting. It employs a time‑aware patch encoding to embed variable‑length intra‑patch timestamps, a time bias attention mechanism to adjust for temporal misalignment and asynchronous cross‑channel dependencies, and a hybrid causal mask on a decoder‑only Transformer to balance historical context with autoregressive forecasting. The authors also curate VersaTSA, a 30 B‑observation dataset preserving native sampling sparsity, and demonstrate state‑of‑the‑art zero‑shot performance on three IMTS benchmarks while remaining competitive on regular MTS tasks.

By Zhihao Lin, Li Lin, Qi Zhang, Kaiwen Xia, Shuai Wang, Jialin Qiao