arXiv AI

Beyond Numerical Time Series: A Unified Benchmark for Multimodal Forecasting with Heterogeneous Context

arXiv Machine Learning
Aug 19

TiMi: Empower Time Series Transformers with Multimodal Mixture of Experts

The paper introduces TiMi, a framework that enhances time series transformers with a Multimodal Mixture-of-Experts (MMoE) module to incorporate multimodal data, especially textual information, into forecasting. TiMi leverages large language models to generate future inferences that guide predictions, eliminating the need for explicit representation alignment. Experiments show TiMi achieves state‑of‑the‑art performance on sixteen real‑world multimodal forecasting benchmarks, outperforming advanced baselines while maintaining adaptability and interpretability.

By Jiafeng Lin, Yuxuan Wang, Huakun Luo, Jianmin Wang, Zhongyi Pei
arXiv Machine Learning
Aug 19

SCENARIODIFF: A Scenario-level Guidance Framework for Multimodal Time Series Forecasting--Extended Version

SCENARIODIFF is a hierarchical contextual reasoning framework designed for multimodal time series forecasting, especially in event-driven domains. It processes textual context through three agents—Historical Context, Scenario, and Anchor Guidance—to generate structured signals that condition a Multimodal Diffusion Transformer. The framework also employs Anchor Blended Sampling to locally refine forecast trajectories without retraining, and demonstrates superior performance on the Time‑MMD benchmark.

By Tuan-Binh Tran, Dat Nguyen Cong, Duc-Trong Le, Thanh Trung Huynh, Tung Kieu
arXiv Machine Learning
Jul 16

Overcoming the Modality Gap in Context-Aided Forecasting

arXiv:2603. 12451v4 Announce Type: replace Abstract: Context-aided forecasting (CAF) holds promise for integrating domain knowledge and forward-looking information, enabling AI systems to surpass traditional statistical methods.

By Vincent Zhihao Zheng, \'Etienne Marcotte, Arjun Ashok, Andrew Robert Williams, Lijun Sun, Alexandre Drouin, Valentina Zantedeschi
arXiv AI
5d ago

When Does Text Inform? Benchmarking Information-Theoretic Metrics for Multimodal Time-Series Forecasting

The paper introduces a synthetic benchmark for multimodal time‑series forecasting that evaluates how well text annotations contribute to predictions. By generating controlled signals with semantically correct, incorrect, and irrelevant annotations, the authors can precisely measure the true information content. Six mutual‑information estimators (KSG, MINE, InfoNCE, CCA, PID, and V‑information) are tested, all correctly ranking useful annotations and enabling annotation auditing without model training. The benchmark also highlights each estimator’s limitations and validates findings on seven real datasets, providing practical guidelines for metric implementation.

By Emma Andrews, Gianmarco Mengaldo
arXiv Computation and Language
Aug 25

Semantics or Structure? Auditing Text Sensitivity in Multimodal Time-Series Forecasting

The paper investigates whether multimodal time‑series forecasting models actually use the semantic content of accompanying text. By systematically perturbing the text—replacing it with empty, constant, shuffled, or cross‑domain sentences—the authors find that mean squared error changes by less than 0.5 % across several architectures, indicating that text does not drive performance gains. They also show that removing a co‑shipped numeric column restores the reported improvements, suggesting that the models rely on other signals rather than textual semantics.

By Karthik Sridhar, Atharva Gupta, Nishant Pradhan, Murari Mandal, Dhruv Kumar, Saurabh Deshpande
arXiv Machine Learning
6d ago

Bidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models

The paper introduces SolCloudLLM, a large language model–based framework that fuses sky‑image patches with time‑series data through bidirectional multimodal fusion for short‑term solar forecasting. Experiments on the SIRTA and SKIPP'D datasets show that SolCloudLLM outperforms existing baselines, achieving up to a 25.4% reduction in mean squared error, especially under cloudy conditions and in few‑shot scenarios.

By Ken Chen, Maneesha Perera, Wei Wang, Sachith Seneviratne, Hansani Weeratunge, Saman Halgamuge