arXiv AI

Rethinking Multimodal Fusion for Time Series: Text Modalities Need Constrained Fusion

arXiv:2603. 22372v2 Announce Type: replace-cross Abstract: Recent advances in multimodal learning have motivated the integration of auxiliary modalities such as text or vision into time series (TS) forecasting.

arXiv Machine Learning
23h ago

TiMi: Empower Time Series Transformers with Multimodal Mixture of Experts

arXiv:2602. 21693v2 Announce Type: replace Abstract: Multimodal time series forecasting has garnered significant attention for its potential to provide more accurate predictions than traditional single-modality models by leveraging rich information inherent in other modalities.

By Jiafeng Lin, Yuxuan Wang, Huakun Luo, Jianmin Wang, Zhongyi Pei
arXiv AI
Jun 9

VFEM: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion

arXiv:2510. 03244v2 Announce Type: replace-cross Abstract: Large time series foundation models often adopt channel-independent architectures to handle varying data dimensions, but this design ignores crucial cross-channel dependencies.

By Yanlong Wang, Hang Yu, Jian Xu, Fei Ma, Hongkang Zhang, Tongtong Feng, Zijian Zhang, Shao-Lun Huang, Danny Dongning Sun, Xiao-Ping Zhang