arXiv:2609.15087v1 Announce Type: cross
Abstract: Most time series forecasting benchmarks remain numerical-centric and provide limited support for evaluating contextual information that shapes real-w...
By Peng Chen, Zhihao Zhuang, Hongzhou Chen, Junhao Huang, Aiping Yang, Mengsen Wu, Yiding Liu, Xilin Dai, Zewei Dong
arXiv:2603. 22372v2 Announce Type: replace-cross Abstract: Recent advances in multimodal learning have motivated the integration of auxiliary modalities such as text or vision into time series (TS) forecasting.
By Seunghan Lee, Jun Seo, Jaehoon Lee, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, SoonYoung Lee, Wonbin Ahn
arXiv:2609.24156v1 Announce Type: cross
Abstract: Most existing time series forecasting methods rely solely on numerical observations, overlooking rich contextual information from auxiliary texts. Re...
By Jiayi Liang, Xiaotian Gu, Xinyu Xie, Yuanbin Wu, Xiaoling Wang
arXiv:2607. 06973v1 Announce Type: new Abstract: We introduce a new context-enriched, multimodal time series forecasting benchmark, TimesX.
By Haoxin Liu, Yichen Zhou, Rajat Sen, B. Aditya Prakash, Abhimanyu Das
SCENARIODIFF is a hierarchical contextual reasoning framework designed for multimodal time series forecasting, especially in event-driven domains. It processes textual context through three agents—Historical Context, Scenario, and Anchor Guidance—to generate structured signals that condition a Multimodal Diffusion Transformer. The framework also employs Anchor Blended Sampling to locally refine forecast trajectories without retraining, and demonstrates superior performance on the Time‑MMD benchmark.
By Tuan-Binh Tran, Dat Nguyen Cong, Duc-Trong Le, Thanh Trung Huynh, Tung Kieu
The paper introduces Expert Modulation, a novel approach for multi‑modal time series prediction that conditions both expert routing and computation on textual signals, thereby providing direct cross‑modal control over expert behavior. Unlike previous methods that rely on token‑level fusion, this mechanism avoids mixing temporal patches with language tokens in a shared embedding space, which can be problematic when high‑quality time‑text pairs are scarce or when time series characteristics vary widely. Experiments and theoretical analysis demonstrate that Expert Modulation yields strong improvements over existing multi‑modal forecasting techniques.
By Lige Zhang, Ali Maatouk, Jialin Chen, Karthik Charan Konduri, Leandros Tassiulas, Rex Ying