arXiv Machine Learning

A Mechanistic Analysis of Transformers for Dynamical Systems

arXiv:2512. 21113v2 Announce Type: replace Abstract: Transformers are increasingly adopted for modeling and forecasting time-series, yet their internal mechanisms remain poorly understood from a dynamical systems perspective.

arXiv Machine Learning
Jun 10

One Step Closer to Ground Truth: A Multi-Scale Residual-Aware Representation Learning Pipeline for Predicting Time Series Data

arXiv:2606. 10678v1 Announce Type: new Abstract: Transformer-based models have emerged as leading paradigms in time-series forecasting in recent years, employing self-attention mechanisms to capture long-range dependencies.

By Amrijit Biswas, Mustafa Kamal, Robin Krambroeckers, M. M. Lutfe Elahi, Sifat Momen, Nabeel Mohammed, Shafin Rahman
arXiv Machine Learning
Jun 2

FAiT: Frequency-Aware Inverted Transformer for Multivariate Time Series Forecasting

arXiv:2606. 01306v1 Announce Type: new Abstract: While Transformer-based architectures have established themselves as a dominant paradigm in Multivariate Time Series Forecasting (MTSF), their core self-attention mechanism inherently functions as a low-pass filter, systematically smoothing out high-frequency signals vital for sharp local changes.

By Peng He, Yao Liu, Yanglei Gan, Run Lin, Yuxiang Cai, Qiao Liu
arXiv Machine Learning
Sep 18

SETTer: Sparse-Encoder Transformer for Long-term Multivariate Time Series Forecasting

SETTer is a transformer-based model designed for long‑term multivariate time‑series forecasting. It introduces decoupled self‑attention and hybrid masking to better handle high dimensionality and complex relationships, while adding explainable structures to highlight discriminative patterns. Experiments on real‑world benchmarks show that SETTer outperforms state‑of‑the‑art models in 88% of scenarios.

By Abraham Ezema, Chijioke Eze, Ferdinanda Ponci, Antonello Monti
arXiv AI
Sep 17

The Attention Within: Consensus Dynamics in Selective State Space Models

Selective state space models (SSMs) use a recurrence to mix token information, a process analogous to attention in transformers. By modeling token evolution as an ordinary differential equation and applying input‑to‑state stability, the study proves that SSMs exhibit local exponential stability of consensus equilibria and delineates their domain of attraction for time‑varying weight matrices. Experiments on a pretrained Mamba‑2 model reveal that the output gate controls the degree of consensus, preventing tokens from fully converging.

By Jo\~ao Pedro Silvestre, \'Alvaro Rodr\'iguez Abella, Paulo Tabuada
Hugging Face Trending Papers
Sep 17

SETTer: Sparse-Encoder Transformer for Long-term Multivariate Time Series Forecasting

SETTer is a transformer-based model designed for long‑term multivariate time‑series forecasting. It introduces decoupled self‑attention and hybrid masking to better capture short‑ and long‑term patterns across time and channel dimensions, while adding simple explainable structures to highlight discriminative patterns. Experiments on real‑world benchmarks show that a single‑layer SETTer outperforms state‑of‑the‑art models in 88% of scenarios.

arXiv Machine Learning
Jun 8

Trio: Learning Time-Series Forecasting with Temporal-Spatial-Sample Attention and Structural Causal Priors

arXiv:2606. 07291v1 Announce Type: new Abstract: Multivariate time-series forecasting requires models to reason over temporal dynamics, cross-variable dependencies, and historical input-output correspondences.

By Tao Chen, Yexu Zhou, Zhi Gong, Hengwei He, Hongda Li, Zhewei Chen, Dongjing Wang, Xin Zhang, Decheng Liu, Chunlei Peng, Zheng Chen, Wenyue Ding
arXiv Machine Learning
Sep 22

A Hybrid Attention Model Learning Unified Time-aware Patch Representation for Irregular Multivariate Time Series Forecasting

The paper introduces a hybrid attention model that learns a unified time‑aware patch representation for irregular multivariate time series (IMTS) forecasting. It employs a time‑aware patch encoding to embed variable‑length intra‑patch timestamps, a time bias attention mechanism to adjust for temporal misalignment and asynchronous cross‑channel dependencies, and a hybrid causal mask on a decoder‑only Transformer to balance historical context with autoregressive forecasting. The authors also curate VersaTSA, a 30 B‑observation dataset preserving native sampling sparsity, and demonstrate state‑of‑the‑art zero‑shot performance on three IMTS benchmarks while remaining competitive on regular MTS tasks.

By Zhihao Lin, Li Lin, Qi Zhang, Kaiwen Xia, Shuai Wang, Jialin Qiao