The paper introduces TiMi, a framework that enhances time series transformers with a Multimodal Mixture-of-Experts (MMoE) module to incorporate multimodal data, especially textual information, into forecasting. TiMi leverages large language models to generate future inferences that guide predictions, eliminating the need for explicit representation alignment. Experiments show TiMi achieves state‑of‑the‑art performance on sixteen real‑world multimodal forecasting benchmarks, outperforming advanced baselines while maintaining adaptability and interpretability.
By Jiafeng Lin, Yuxuan Wang, Huakun Luo, Jianmin Wang, Zhongyi Pei
arXiv:2607. 06973v1 Announce Type: new Abstract: We introduce a new context-enriched, multimodal time series forecasting benchmark, TimesX.
By Haoxin Liu, Yichen Zhou, Rajat Sen, B. Aditya Prakash, Abhimanyu Das
SCENARIODIFF is a hierarchical contextual reasoning framework designed for multimodal time series forecasting, especially in event-driven domains. It processes textual context through three agents—Historical Context, Scenario, and Anchor Guidance—to generate structured signals that condition a Multimodal Diffusion Transformer. The framework also employs Anchor Blended Sampling to locally refine forecast trajectories without retraining, and demonstrates superior performance on the Time‑MMD benchmark.
By Tuan-Binh Tran, Dat Nguyen Cong, Duc-Trong Le, Thanh Trung Huynh, Tung Kieu
Textual context such as news, reports, and logs can provide valuable signals for time series forecasting, especially when future dynamics are driven by external events that are not yet visible in hist...
arXiv:2603. 12451v4 Announce Type: replace Abstract: Context-aided forecasting (CAF) holds promise for integrating domain knowledge and forward-looking information, enabling AI systems to surpass traditional statistical methods.
By Vincent Zhihao Zheng, \'Etienne Marcotte, Arjun Ashok, Andrew Robert Williams, Lijun Sun, Alexandre Drouin, Valentina Zantedeschi
arXiv:2602. 01588v3 Announce Type: replace-cross Abstract: Multimodal time series forecasting is crucial in real-world applications, where decisions depend on both numerical data and contextual signals.
By Huu Hiep Nguyen, Minh Hoang Nguyen, Dung Nguyen, Hung Le
arXiv:2603. 22372v2 Announce Type: replace-cross Abstract: Recent advances in multimodal learning have motivated the integration of auxiliary modalities such as text or vision into time series (TS) forecasting.
By Seunghan Lee, Jun Seo, Jaehoon Lee, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, SoonYoung Lee, Wonbin Ahn
The paper introduces a synthetic benchmark for multimodal time‑series forecasting that evaluates how well text annotations contribute to predictions. By generating controlled signals with semantically correct, incorrect, and irrelevant annotations, the authors can precisely measure the true information content. Six mutual‑information estimators (KSG, MINE, InfoNCE, CCA, PID, and V‑information) are tested, all correctly ranking useful annotations and enabling annotation auditing without model training. The benchmark also highlights each estimator’s limitations and validates findings on seven real datasets, providing practical guidelines for metric implementation.
By Emma Andrews, Gianmarco Mengaldo
The paper investigates whether multimodal time‑series forecasting models actually use the semantic content of accompanying text. By systematically perturbing the text—replacing it with empty, constant, shuffled, or cross‑domain sentences—the authors find that mean squared error changes by less than 0.5 % across several architectures, indicating that text does not drive performance gains. They also show that removing a co‑shipped numeric column restores the reported improvements, suggesting that the models rely on other signals rather than textual semantics.
By Karthik Sridhar, Atharva Gupta, Nishant Pradhan, Murari Mandal, Dhruv Kumar, Saurabh Deshpande
The paper introduces SolCloudLLM, a large language model–based framework that fuses sky‑image patches with time‑series data through bidirectional multimodal fusion for short‑term solar forecasting. Experiments on the SIRTA and SKIPP'D datasets show that SolCloudLLM outperforms existing baselines, achieving up to a 25.4% reduction in mean squared error, especially under cloudy conditions and in few‑shot scenarios.
By Ken Chen, Maneesha Perera, Wei Wang, Sachith Seneviratne, Hansani Weeratunge, Saman Halgamuge
arXiv:2603. 05997v2 Announce Type: replace-cross Abstract: Irregularly sampled time series (ISTS) are widespread in real-world scenarios, exhibiting asynchronous observations on uneven time intervals across diverse variables.
By Zhi Lei, Chenxi Liu, Hao Miao, Wanghui Qiu, Bin Yang, Chenjuan Guo
arXiv:2606. 06285v1 Announce Type: new Abstract: Time series foundation models (TS-FMs) aim to learn generalizable temporal representations that can be adapted to a wide range of downstream tasks.
By Ziwen Kan, Yishuo Chen, Kecheng Li, Andrew Wen, Xiaomeng Wang, Liwei Wang, Jihao Duan, Song Wang, Hongfang Liu, Tianlong Chen