arXiv:2609.24441v1 Announce Type: new
Abstract: Multivariate time-series forecasting is essential to many real-world applications. Recent large vision models (LVMs) offer a promising paradigm by tran...
By Xinying Cai, Junkai Lu, Yuhan Zhu, Xiaoyun Yu, Xiangfei Qiu, Jilin Hu
Multivariate time-series forecasting is essential to many real-world applications. Recent large vision models (LVMs) offer a promising paradigm by transferring cross-domain visual priors to time-serie...
arXiv:2603. 05997v2 Announce Type: replace-cross Abstract: Irregularly sampled time series (ISTS) are widespread in real-world scenarios, exhibiting asynchronous observations on uneven time intervals across diverse variables.
By Zhi Lei, Chenxi Liu, Hao Miao, Wanghui Qiu, Bin Yang, Chenjuan Guo
arXiv:2603. 22372v2 Announce Type: replace-cross Abstract: Recent advances in multimodal learning have motivated the integration of auxiliary modalities such as text or vision into time series (TS) forecasting.
By Seunghan Lee, Jun Seo, Jaehoon Lee, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, SoonYoung Lee, Wonbin Ahn
arXiv:2602. 01588v3 Announce Type: replace-cross Abstract: Multimodal time series forecasting is crucial in real-world applications, where decisions depend on both numerical data and contextual signals.
By Huu Hiep Nguyen, Minh Hoang Nguyen, Dung Nguyen, Hung Le
The paper introduces ICI-Time, a framework that casts time series forecasting as a visual inpainting problem. By converting series into area‑chart images, it enables pre‑trained vision transformers to perform forecasting through in‑context learning without fine‑tuning or new temporal architectures. Experiments on epidemiology, meteorology, and power systems show competitive performance and strong adaptability in low‑data scenarios.
By Thang Nguyen, Dung Nguyen, Romero Morais, Truyen Tran
We propose ICI-Time, a novel framework that reframes time series forecasting as a visual inpainting task, leveraging the generalisation power of large vision models (LVMs). Unlike methods that require...
WinoTS introduces a wavelet‑based self‑distillation framework for time‑series models that uses time‑frequency augmentations to create multi‑scale structural views, avoiding distortion of signal dynamics. The method outperforms state‑of‑the‑art baselines in long‑term forecasting, cross‑domain zero‑shot transfer, and unsupervised anomaly detection, and linear probing on frozen representations often beats fully supervised training from scratch. Ablation studies show WinoTS is architecture‑agnostic and demonstrates that time‑frequency transformations offer a principled alternative to vision‑style spatial augmentations.
By Noam Major, Kathy Razmadze, Yoli Shavit
WorldTS is a new forecasting framework that models latent dynamics conditioned on multimodal covariates to improve time‑series prediction. It uses a two‑stage training process: first learning latent state dynamics from historical data and covariates, then training a decoder to map predicted latent states back to future observations. Experiments on 21 real‑world datasets demonstrate the effectiveness of this approach.
By Yuhan Zhu, Xiangfei Qiu, Hanyin Cheng, Wangmeng Shen, Chenjuan Guo, Bin Yang, Jilin Hu, Christian S. Jensen
SAGE is an end‑to‑end CLIP‑based framework that augments vision‑language time‑series forecasting by jointly modeling temporal, cross‑variable, textual, and visual information. It processes frequency‑enhanced patches and variable tokens through a CLIP text encoder, while gated residual paths inject variable‑specific descriptions and statistical descriptors. A frozen CLIP vision encoder aligns rendered series with temporal representations via a training‑only contrastive objective, enabling multimodal alignment and variable‑level knowledge without using an LLM during inference.
arXiv:2607. 19404v1 Announce Type: cross Abstract: Multivariate time series encode structural patterns that unfold across multiple temporal scales, yet most forecasting backbones treat learned representations as transient byproducts of prediction, leaving the organizational geometry of these patterns underexploited.
By Xingsheng Chen, Deyu Yi, Siu-Ming Yiu
The paper introduces m-WCN, an end‑to‑end deep learning framework that neuralizes multi‑wavelet decomposition to jointly extract temporal patterns and frequency components from time series. Two task‑specific architectures built on m‑WCN—TFBC for classification and FTB for forecasting—are shown to outperform baseline models on 64 UCR datasets and seven forecasting benchmarks, achieving average improvements of nearly 20% in both tasks. The approach leverages trainable convolutional operators and orthogonality constraints to produce interpretable multi‑resolution representations.
By Xiaohan Jiang, Jingyuan Wang, Jiahao Ji, Yongyao Wang, Chen Yang, Junjie Wu