Hugging Face Trending Papers

MUSE: Dependency-Aware Adaptation of a Frozen Vision Backbone for Multivariate Time Series Forecasting

arXiv AI
Jun 9

VFEM: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion

arXiv:2510. 03244v2 Announce Type: replace-cross Abstract: Large time series foundation models often adopt channel-independent architectures to handle varying data dimensions, but this design ignores crucial cross-channel dependencies.

By Yanlong Wang, Hang Yu, Jian Xu, Fei Ma, Hongkang Zhang, Tongtong Feng, Zijian Zhang, Shao-Lun Huang, Danny Dongning Sun, Xiao-Ping Zhang
arXiv AI
Aug 26

In-Context Inpainting for Time Series Forecasting

The paper introduces ICI-Time, a framework that casts time series forecasting as a visual inpainting problem. By converting series into area‑chart images, it enables pre‑trained vision transformers to perform forecasting through in‑context learning without fine‑tuning or new temporal architectures. Experiments on epidemiology, meteorology, and power systems show competitive performance and strong adaptability in low‑data scenarios.

By Thang Nguyen, Dung Nguyen, Romero Morais, Truyen Tran
Hugging Face Trending Papers
Aug 27

SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting

SAGE is an end‑to‑end CLIP‑based framework that augments vision‑language time‑series forecasting by jointly modeling temporal, cross‑variable, textual, and visual information. It processes frequency‑enhanced patches and variable tokens through a CLIP text encoder, while gated residual paths inject variable‑specific descriptions and statistical descriptors. A frozen CLIP vision encoder aligns rendered series with temporal representations via a training‑only contrastive objective, enabling multimodal alignment and variable‑level knowledge without using an LLM during inference.

arXiv Machine Learning
Aug 28

SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting

SAGE is a CLIP-based framework that augments vision‑language time series forecasting by incorporating variable‑specific semantic and statistical information. It processes frequency‑enhanced patches and variable tokens through a CLIP text encoder, while a frozen CLIP vision encoder aligns rendered series with temporal representations via a contrastive objective. The approach achieves state‑of‑the‑art accuracy on eight long‑term benchmarks and M4, with ablations showing complementary gains from multimodal alignment and variable‑level knowledge.

By Haizhao Fan, Xinyi Le
arXiv AI
1d ago

WinoTS: Wavelet-based Self-Distillation for Time Series Models

WinoTS introduces a wavelet‑based self‑distillation framework for time‑series models that uses time‑frequency augmentations to create multi‑scale structural views, avoiding distortion of signal dynamics. The method outperforms state‑of‑the‑art baselines in long‑term forecasting, cross‑domain zero‑shot transfer, and unsupervised anomaly detection, and linear probing on frozen representations often beats fully supervised training from scratch. Ablation studies show WinoTS is architecture‑agnostic and demonstrates that time‑frequency transformations offer a principled alternative to vision‑style spatial augmentations.

By Noam Major, Kathy Razmadze, Yoli Shavit
arXiv AI
Sep 25

TimeBraid: Unifying Time Series and Language for Understanding and Forecasting

TimeBraid is a family of unified models that combine pretrained language models with pretrained time‑series foundation models using interleaved global residual attention layers. The models inherit instruction following, reasoning, and continuous‑signal perception, fusing both modalities into a shared representation space for understanding and generation. The design focuses on aligning representation spaces, grounding language in temporal structure, balancing understanding with generation, and maintaining stable joint optimization, supported by 2.2 M curated series‑text pairs and 4.9 M instruction‑tuning samples. Across diverse benchmarks, TimeBraid competes with larger general‑purpose and task‑specific models.

By Xinyue Wang, Jiacheng Pang, Kun Zhou, Kexin Zhang, Defu Cao, Fan Feng, Faisal, Songyao Jin, Yan Liu, Biwei Huang