arXiv AI

Signed Dual Attention: Capturing Signed Dependencies in Time Series Forecasting

arXiv:2606. 04833v1 Announce Type: cross Abstract: Initially developed for natural language processing, Transformer architectures and attention mechanisms are now central to a wide range of deep learning models, including applications in time series forecasting.

arXiv AI
Sep 25

TimeBraid: Unifying Time Series and Language for Understanding and Forecasting

TimeBraid is a family of unified models that combine pretrained language models with pretrained time‑series foundation models using interleaved global residual attention layers. The models inherit instruction following, reasoning, and continuous‑signal perception, fusing both modalities into a shared representation space for understanding and generation. The design focuses on aligning representation spaces, grounding language in temporal structure, balancing understanding with generation, and maintaining stable joint optimization, supported by 2.2 M curated series‑text pairs and 4.9 M instruction‑tuning samples. Across diverse benchmarks, TimeBraid competes with larger general‑purpose and task‑specific models.

By Xinyue Wang, Jiacheng Pang, Kun Zhou, Kexin Zhang, Defu Cao, Fan Feng, Faisal, Songyao Jin, Yan Liu, Biwei Huang
Hugging Face Trending Papers
Aug 6

TS-RAG: Retrieval Augmented Generation for Time Series Forecasting

While deep learning models, particularly transformer-based architectures, have shown impressive performance in time series forecasting, the application of retrieval-augmented generation (RAG) in this domain remains limited. Since RAG has proven effective in enhancing the capabilities of large language models by incorporating relevant external information, retrieving similar time series sequences as references might also improve accuracy in time series forecasting tasks.

arXiv Machine Learning
Sep 25

Beyond Pairwise Attention: Higher-Order Modular Attention for Efficient Sequence Learning

The paper introduces Higher-Order Modular Attention (HOMA), a new attention mechanism that combines standard pairwise self‑attention with an explicit triadic attention pathway. HOMA uses overlapping blocks, local windows, and a low‑rank projection to make triadic interactions tractable. Experiments on controlled PARITY and MATCH3 tasks, as well as TAPE benchmarks, show that HOMA matches or outperforms matched pairwise and purely triadic baselines, especially when dependencies extend beyond triadic order, and it often converges faster and uses parameters more efficiently.

By Shirin Amiraslani, Xin Gao
arXiv AI
Sep 10

Fewer yet critical: Reducing Redundant Token Dependencies for Transformer-based Time Series Forecasting

The paper introduces a token dependency selection strategy for Transformer-based time series forecasting. By jointly applying an attention entropy constraint and a prediction error constraint, the method identifies fewer but more critical inter-token dependencies, reducing the influence of redundant dependencies that can hurt generalization. Experiments on multiple datasets show that this approach improves forecasting performance across various Transformer models.

By Jianqi Zhang, Yuchan Liu, Zeen Song, Yuefei Li, Fanjiang Xu
arXiv Machine Learning
Sep 29

Progressive Memory Transformer: Memory-Aware Attention for Time-Series

The paper introduces the Progressive Memory Transformer (PMT), a transformer variant that adds writable, window‑aligned memory to expose mid‑range representations alongside token and sequence‑level outputs. PMT is trained with a hierarchical learning framework that applies separate objectives at local, mid‑range, and global scales, encouraging the model to capture fine‑grained variation, window‑level motifs, and overall sequence agreement. Experiments on seven UCR/UEA/UCI classification datasets, a cue‑retention probe, and forecasting tasks show that PMT achieves strong low‑label classification performance, competitive multi‑horizon forecasting, and evidence that its memory states encode mid‑range motifs.

By Tord Sture Stangeland, Andreas K\"ohler, Steffen M{\ae}land, Ad\'in Ram\'ires Rivera
arXiv AI
Jul 1

Mantis: Lightweight Foundation Model for Time Series Classification

arXiv:2502. 15637v2 Announce Type: replace-cross Abstract: While foundation models have revolutionized various domains, their application to time series classification remains rather under-explored, with existing literature predominantly focused on forecasting.

By Vasilii Feofanov, Songkang Wen, Shifeng Xie, Simon Roschmann, Marius Alonso, Hongbo Guo, Romain Ilbert, Malik Tiomoko, Quentin Bouniot, Zeynep Akata, Lujia Pan, Jianfeng Zhang, Ievgen Redko
arXiv AI
Aug 11

Linearized 2-Simplicial Attention

arXiv:2608. 09307v1 Announce Type: new Abstract: We present a linearized form of 2-simplicial attention by rewriting the trilinear score as an inner product between a composite query and a key, so that the sum over one token axis takes the same form as ordinary softmax attention.

By Aritra Das, Dhruman Gupta, Debayan Gupta