arXiv AI

TimeOmni-VL: Unified Models for Time Series Understanding and Generation

arXiv:2602. 17149v2 Announce Type: replace-cross Abstract: Recent time series modeling faces a sharp divide between numerical generation and semantic understanding, with research showing that generation models often rely on superficial pattern matching, while understanding-oriented models struggle with high-fidelity numerical output.

arXiv AI
Jun 6

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models

arXiv:2606. 05702v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have significantly enhanced their ability to interpret complex visual semantics, yet their capacity for chronological reasoning remains under-explored.

By Haoyu Zhou, Qing Qing, Caichong Li, Qixin Zhang, Yongcheng Jing, Ziqi Xu, Juncheng Hu, Xikun Zhang, Renqiang Luo
arXiv Computer Vision
Aug 25

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

SynerMedGen is a unified framework that aligns medical multimodal understanding with generation tasks through task alignment. It introduces three generation‑aligned understanding tasks and a two‑stage training strategy that transfers representations learned during understanding to medical image synthesis. The model achieves strong zero‑shot performance on 22 synthesis tasks and outperforms state‑of‑the‑art specialized and unified models when combined with generation training, supported by a new 1M‑sample SynerMed dataset.

By Weiren Zhao, Yi Dong, Cheng Chen
arXiv AI
Sep 25

TimeBraid: Unifying Time Series and Language for Understanding and Forecasting

TimeBraid is a family of unified models that combine pretrained language models with pretrained time‑series foundation models using interleaved global residual attention layers. The models inherit instruction following, reasoning, and continuous‑signal perception, fusing both modalities into a shared representation space for understanding and generation. The design focuses on aligning representation spaces, grounding language in temporal structure, balancing understanding with generation, and maintaining stable joint optimization, supported by 2.2 M curated series‑text pairs and 4.9 M instruction‑tuning samples. Across diverse benchmarks, TimeBraid competes with larger general‑purpose and task‑specific models.

By Xinyue Wang, Jiacheng Pang, Kun Zhou, Kexin Zhang, Defu Cao, Fan Feng, Faisal, Songyao Jin, Yan Liu, Biwei Huang
arXiv Machine Learning
Sep 24

ChronoSteer: Bridging Large Language Model and Time Series Foundation Model via Synthetic Cross-Modal Alignment Dataset

ChronoSteer is a decoupled agentic framework that bridges large language models and time series foundation models by learning cross‑modal alignment from synthetic paired supervision. It converts textual events into revision instructions that steer a frozen time‑series model, discretizes these instructions into a compact codebook to reduce semantic divergence, and then refines the predictions with a two‑stage training strategy. The authors also release a leakage‑controlled multimodal benchmark and report a 25.8% improvement in zero‑shot prediction accuracy over the unimodal backbone.

By Chengsen Wang, Qi Qi, Zhongwen Rao, Lujia Pan, Jingyu Wang
arXiv Machine Learning
Aug 19

TiMi: Empower Time Series Transformers with Multimodal Mixture of Experts

The paper introduces TiMi, a framework that enhances time series transformers with a Multimodal Mixture-of-Experts (MMoE) module to incorporate multimodal data, especially textual information, into forecasting. TiMi leverages large language models to generate future inferences that guide predictions, eliminating the need for explicit representation alignment. Experiments show TiMi achieves state‑of‑the‑art performance on sixteen real‑world multimodal forecasting benchmarks, outperforming advanced baselines while maintaining adaptability and interpretability.

By Jiafeng Lin, Yuxuan Wang, Huakun Luo, Jianmin Wang, Zhongyi Pei