arXiv AI

VFEM: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion

arXiv:2510. 03244v2 Announce Type: replace-cross Abstract: Large time series foundation models often adopt channel-independent architectures to handle varying data dimensions, but this design ignores crucial cross-channel dependencies.

arXiv AI
Aug 26

In-Context Inpainting for Time Series Forecasting

The paper introduces ICI-Time, a framework that casts time series forecasting as a visual inpainting problem. By converting series into area‑chart images, it enables pre‑trained vision transformers to perform forecasting through in‑context learning without fine‑tuning or new temporal architectures. Experiments on epidemiology, meteorology, and power systems show competitive performance and strong adaptability in low‑data scenarios.

By Thang Nguyen, Dung Nguyen, Romero Morais, Truyen Tran
arXiv AI
3d ago

WinoTS: Wavelet-based Self-Distillation for Time Series Models

WinoTS introduces a wavelet‑based self‑distillation framework for time‑series models that uses time‑frequency augmentations to create multi‑scale structural views, avoiding distortion of signal dynamics. The method outperforms state‑of‑the‑art baselines in long‑term forecasting, cross‑domain zero‑shot transfer, and unsupervised anomaly detection, and linear probing on frozen representations often beats fully supervised training from scratch. Ablation studies show WinoTS is architecture‑agnostic and demonstrates that time‑frequency transformations offer a principled alternative to vision‑style spatial augmentations.

By Noam Major, Kathy Razmadze, Yoli Shavit
arXiv Machine Learning
5d ago

WorldTS: World Modeling for Multimodal Covariate-aware Time Series Forecasting

WorldTS is a new forecasting framework that models latent dynamics conditioned on multimodal covariates to improve time‑series prediction. It uses a two‑stage training process: first learning latent state dynamics from historical data and covariates, then training a decoder to map predicted latent states back to future observations. Experiments on 21 real‑world datasets demonstrate the effectiveness of this approach.

By Yuhan Zhu, Xiangfei Qiu, Hanyin Cheng, Wangmeng Shen, Chenjuan Guo, Bin Yang, Jilin Hu, Christian S. Jensen
Hugging Face Trending Papers
Aug 27

SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting

SAGE is an end‑to‑end CLIP‑based framework that augments vision‑language time‑series forecasting by jointly modeling temporal, cross‑variable, textual, and visual information. It processes frequency‑enhanced patches and variable tokens through a CLIP text encoder, while gated residual paths inject variable‑specific descriptions and statistical descriptors. A frozen CLIP vision encoder aligns rendered series with temporal representations via a training‑only contrastive objective, enabling multimodal alignment and variable‑level knowledge without using an LLM during inference.

arXiv AI
Jul 23

Structured Latent Space Modeling over Multi-Scale Temporal Patches for Multivariate Time Series Forecasting

arXiv:2607. 19404v1 Announce Type: cross Abstract: Multivariate time series encode structural patterns that unfold across multiple temporal scales, yet most forecasting backbones treat learned representations as transient byproducts of prediction, leaving the organizational geometry of these patterns underexploited.

By Xingsheng Chen, Deyu Yi, Siu-Ming Yiu
arXiv AI
Sep 25

Neuralized Multi-Wavelet Decomposition for Time Series Classification and Forecasting

The paper introduces m-WCN, an end‑to‑end deep learning framework that neuralizes multi‑wavelet decomposition to jointly extract temporal patterns and frequency components from time series. Two task‑specific architectures built on m‑WCN—TFBC for classification and FTB for forecasting—are shown to outperform baseline models on 64 UCR datasets and seven forecasting benchmarks, achieving average improvements of nearly 20% in both tasks. The approach leverages trainable convolutional operators and orthogonality constraints to produce interpretable multi‑resolution representations.

By Xiaohan Jiang, Jingyuan Wang, Jiahao Ji, Yongyao Wang, Chen Yang, Junjie Wu