arXiv:2603. 05997v2 Announce Type: replace-cross Abstract: Irregularly sampled time series (ISTS) are widespread in real-world scenarios, exhibiting asynchronous observations on uneven time intervals across diverse variables.
By Zhi Lei, Chenxi Liu, Hao Miao, Wanghui Qiu, Bin Yang, Chenjuan Guo
The paper introduces TiMi, a framework that enhances time series transformers with a Multimodal Mixture-of-Experts (MMoE) module to incorporate multimodal data, especially textual information, into forecasting. TiMi leverages large language models to generate future inferences that guide predictions, eliminating the need for explicit representation alignment. Experiments show TiMi achieves state‑of‑the‑art performance on sixteen real‑world multimodal forecasting benchmarks, outperforming advanced baselines while maintaining adaptability and interpretability.
By Jiafeng Lin, Yuxuan Wang, Huakun Luo, Jianmin Wang, Zhongyi Pei
arXiv:2601. 14968v2 Announce Type: replace-cross Abstract: Most existing time series classification methods adopt a discriminative paradigm that maps input sequences directly to one-hot encoded class labels.
By Mingyue Cheng, Xiaoyu Tao, Huajian Zhang, Qi Liu, Zhiding Liu, Yucong Luo, Yiheng Chen, Enhong Chen
arXiv:2602. 01588v3 Announce Type: replace-cross Abstract: Multimodal time series forecasting is crucial in real-world applications, where decisions depend on both numerical data and contextual signals.
By Huu Hiep Nguyen, Minh Hoang Nguyen, Dung Nguyen, Hung Le
arXiv:2603. 22372v2 Announce Type: replace-cross Abstract: Recent advances in multimodal learning have motivated the integration of auxiliary modalities such as text or vision into time series (TS) forecasting.
By Seunghan Lee, Jun Seo, Jaehoon Lee, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, SoonYoung Lee, Wonbin Ahn
arXiv:2608.28605v1 Announce Type: new
Abstract: Recent advancements in multimodal learning for medical time series (MedTS) classification highlight the benefits of integrating complementary modalitie...
By Jiexia Ye, Jia Li, Fugee Tsung
arXiv:2606. 06285v1 Announce Type: new Abstract: Time series foundation models (TS-FMs) aim to learn generalizable temporal representations that can be adapted to a wide range of downstream tasks.
By Ziwen Kan, Yishuo Chen, Kecheng Li, Andrew Wen, Xiaomeng Wang, Liwei Wang, Jihao Duan, Song Wang, Hongfang Liu, Tianlong Chen
ChronoSteer is a decoupled agentic framework that bridges large language models and time series foundation models by learning cross‑modal alignment from synthetic paired supervision. It converts textual events into revision instructions that steer a frozen time‑series model, discretizes these instructions into a compact codebook to reduce semantic divergence, and then refines the predictions with a two‑stage training strategy. The authors also release a leakage‑controlled multimodal benchmark and report a 25.8% improvement in zero‑shot prediction accuracy over the unimodal backbone.
By Chengsen Wang, Qi Qi, Zhongwen Rao, Lujia Pan, Jingyu Wang
The paper explores remaining useful life (RUL) estimation using multimodal large language models (MLLMs) that are grounded through time‑series retrieval. It proposes a framework that retrieves historically similar degradation segments, combines them with the test trajectory into a visual comparison artifact, and processes this via a structured multimodal prompt. Experiments on the FD001 partition of the C‑MAPSS benchmark show that retrieval consistently improves RUL prediction, with greater benefits for larger MLLMs, while also revealing current limitations in practical prognostics and health management (PHM) settings.
By Valeriu Dimidov, Rapha\"el Frank
arXiv:2608. 19218v1 Announce Type: cross Abstract: Large language models (LLMs) and agentic AI systems are increasingly being explored for domain-specific maintenance and prognostics tasks, raising the question of whether they can effectively support prognostics and health management (PHM).
By Valeriu Dimidov, Rapha\"el Frank
The paper introduces Expert Modulation, a novel approach for multi‑modal time series prediction that conditions both expert routing and computation on textual signals, thereby providing direct cross‑modal control over expert behavior. Unlike previous methods that rely on token‑level fusion, this mechanism avoids mixing temporal patches with language tokens in a shared embedding space, which can be problematic when high‑quality time‑text pairs are scarce or when time series characteristics vary widely. Experiments and theoretical analysis demonstrate that Expert Modulation yields strong improvements over existing multi‑modal forecasting techniques.
By Lige Zhang, Ali Maatouk, Jialin Chen, Karthik Charan Konduri, Leandros Tassiulas, Rex Ying
The paper introduces a generalized multimodal foundation model that can handle arbitrary combinations of modalities and prediction tasks. It trains on large-scale synthetic multimodal datasets with diverse causal structures to learn transferable multimodal correlations. Experiments on 18 real-world datasets across 12 modalities and 11 tasks show competitive performance compared to specialized models without task-specific adaptation.
By Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, Changqing Zhang