arXiv Machine Learning By Jiafeng Lin, Yuxuan Wang, Huakun Luo, Jianmin Wang, Zhongyi Pei

TiMi: Empower Time Series Transformers with Multimodal Mixture of Experts

Read the original on arXiv Machine Learning →

The paper introduces TiMi, a framework that enhances time series transformers with a Multimodal Mixture-of-Experts (MMoE) module to incorporate multimodal data, especially textual information, into forecasting. TiMi leverages large language models to generate future inferences that guide predictions, eliminating the need for explicit representation alignment. Experiments show TiMi achieves state‑of‑the‑art performance on sixteen real‑world multimodal forecasting benchmarks, outperforming advanced baselines while maintaining adaptability and interpretability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 19

SCENARIODIFF: A Scenario-level Guidance Framework for Multimodal Time Series Forecasting--Extended Version

SCENARIODIFF is a hierarchical contextual reasoning framework designed for multimodal time series forecasting, especially in event-driven domains. It processes textual context through three agents—Historical Context, Scenario, and Anchor Guidance—to generate structured signals that condition a Multimodal Diffusion Transformer. The framework also employs Anchor Blended Sampling to locally refine forecast trajectories without retraining, and demonstrates superior performance on the Time‑MMD benchmark.

By Tuan-Binh Tran, Dat Nguyen Cong, Duc-Trong Le, Thanh Trung Huynh, Tung Kieu
arXiv AI
Sep 7

Multi-Modal Time Series Prediction via Mixture of Modulated Experts

The paper introduces Expert Modulation, a novel approach for multi‑modal time series prediction that conditions both expert routing and computation on textual signals, thereby providing direct cross‑modal control over expert behavior. Unlike previous methods that rely on token‑level fusion, this mechanism avoids mixing temporal patches with language tokens in a shared embedding space, which can be problematic when high‑quality time‑text pairs are scarce or when time series characteristics vary widely. Experiments and theoretical analysis demonstrate that Expert Modulation yields strong improvements over existing multi‑modal forecasting techniques.

By Lige Zhang, Ali Maatouk, Jialin Chen, Karthik Charan Konduri, Leandros Tassiulas, Rex Ying