arXiv AI By Zhi Lei, Chenxi Liu, Hao Miao, Wanghui Qiu, Bin Yang, Chenjuan Guo

MM-ISTS: Cooperating Irregularly Sampled Time Series Forecasting with Multimodal Vision-Text LLMs

Read the original on arXiv AI →

arXiv:2603. 05997v2 Announce Type: replace-cross Abstract: Irregularly sampled time series (ISTS) are widespread in real-world scenarios, exhibiting asynchronous observations on uneven time intervals across diverse variables.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 28

SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting

SAGE is a CLIP-based framework that augments vision‑language time series forecasting by incorporating variable‑specific semantic and statistical information. It processes frequency‑enhanced patches and variable tokens through a CLIP text encoder, while a frozen CLIP vision encoder aligns rendered series with temporal representations via a contrastive objective. The approach achieves state‑of‑the‑art accuracy on eight long‑term benchmarks and M4, with ablations showing complementary gains from multimodal alignment and variable‑level knowledge.

By Haizhao Fan, Xinyi Le
arXiv Machine Learning
Aug 19

TiMi: Empower Time Series Transformers with Multimodal Mixture of Experts

The paper introduces TiMi, a framework that enhances time series transformers with a Multimodal Mixture-of-Experts (MMoE) module to incorporate multimodal data, especially textual information, into forecasting. TiMi leverages large language models to generate future inferences that guide predictions, eliminating the need for explicit representation alignment. Experiments show TiMi achieves state‑of‑the‑art performance on sixteen real‑world multimodal forecasting benchmarks, outperforming advanced baselines while maintaining adaptability and interpretability.

By Jiafeng Lin, Yuxuan Wang, Huakun Luo, Jianmin Wang, Zhongyi Pei
Hugging Face Trending Papers
Aug 27

SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting

SAGE is an end‑to‑end CLIP‑based framework that augments vision‑language time‑series forecasting by jointly modeling temporal, cross‑variable, textual, and visual information. It processes frequency‑enhanced patches and variable tokens through a CLIP text encoder, while gated residual paths inject variable‑specific descriptions and statistical descriptors. A frozen CLIP vision encoder aligns rendered series with temporal representations via a training‑only contrastive objective, enabling multimodal alignment and variable‑level knowledge without using an LLM during inference.

arXiv Machine Learning
Sep 24

ChronoSteer: Bridging Large Language Model and Time Series Foundation Model via Synthetic Cross-Modal Alignment Dataset

ChronoSteer is a decoupled agentic framework that bridges large language models and time series foundation models by learning cross‑modal alignment from synthetic paired supervision. It converts textual events into revision instructions that steer a frozen time‑series model, discretizes these instructions into a compact codebook to reduce semantic divergence, and then refines the predictions with a two‑stage training strategy. The authors also release a leakage‑controlled multimodal benchmark and report a 25.8% improvement in zero‑shot prediction accuracy over the unimodal backbone.

By Chengsen Wang, Qi Qi, Zhongwen Rao, Lujia Pan, Jingyu Wang