MUSE: Dependency-Aware Adaptation of a Frozen Vision Backbone for Multivariate Time Series Forecasting
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
Multivariate time-series forecasting is essential to many real-world applications. Recent large vision models (LVMs) offer a promising paradigm by transferring cross-domain visual priors to time-serie...
arXiv:2510. 03244v2 Announce Type: replace-cross Abstract: Large time series foundation models often adopt channel-independent architectures to handle varying data dimensions, but this design ignores crucial cross-channel dependencies.
The paper introduces ICI-Time, a framework that casts time series forecasting as a visual inpainting problem. By converting series into area‑chart images, it enables pre‑trained vision transformers to perform forecasting through in‑context learning without fine‑tuning or new temporal architectures. Experiments on epidemiology, meteorology, and power systems show competitive performance and strong adaptability in low‑data scenarios.
We propose ICI-Time, a novel framework that reframes time series forecasting as a visual inpainting task, leveraging the generalisation power of large vision models (LVMs). Unlike methods that require...
arXiv:2506.08641v3 Announce Type: replace Abstract: Adapting vision models for time series analysis is compelling, yet all existing approaches are falling short of dedicated time series foundation mo...
SAGE is an end‑to‑end CLIP‑based framework that augments vision‑language time‑series forecasting by jointly modeling temporal, cross‑variable, textual, and visual information. It processes frequency‑enhanced patches and variable tokens through a CLIP text encoder, while gated residual paths inject variable‑specific descriptions and statistical descriptors. A frozen CLIP vision encoder aligns rendered series with temporal representations via a training‑only contrastive objective, enabling multimodal alignment and variable‑level knowledge without using an LLM during inference.