arXiv Machine Learning

A Lightweight Plug-in Gate for Transformer-Based Time-Series Forecasters

The paper introduces a lightweight pre‑encoder gate for Transformer‑based time‑series forecasters, which assigns sigmoid scores to covariate representations before they enter the encoder. The gate is evaluated as a plug‑in for models such as TimeXer, iTransformer, and PatchTST on datasets including ETTm1, ETTm2, Traffic, Energy, and ILI, showing competitive performance and the ability to regulate covariate admission via a usage penalty. Experiments also explore gate placement, initialization, and feature importance using VIF‑informed permutation diagnostics.

arXiv Machine Learning
Jul 20

A Benchmark for Electrical Load Forecasting Across Grid Levels: Time-Series Transformers Outperform Established Methods

arXiv:2607. 15705v1 Announce Type: new Abstract: Accurate load forecasting at multiple grid levels is essential for future smart grids, ranging from aggregated control area forecasts for balancing supply and demand to forecasts of individual end-consumer loads for demand-side management and energy management systems.

By Matthias Hertel, Sebastian P\"utz, Jonathan Kolar, Benjamin Sch\"afer, Ralf Mikut, Veit Hagenmeyer
arXiv Machine Learning
Jun 8

CF-JEPA: Mask-free forward prediction with asymmetric encoder utilization for time-series representation learning

arXiv:2606. 07031v1 Announce Type: new Abstract: Self-supervised learning (SSL) for time-series representation learning is dominated by two paradigms: contrastive methods, which face challenges in constructing positive or negative pairs, and masking-based methods, which disrupt the temporal continuity of time-series signals.

By Jaehoon Lee, Sunghyun Sim
arXiv Machine Learning
Sep 18

SETTer: Sparse-Encoder Transformer for Long-term Multivariate Time Series Forecasting

SETTer is a transformer-based model designed for long‑term multivariate time‑series forecasting. It introduces decoupled self‑attention and hybrid masking to better handle high dimensionality and complex relationships, while adding explainable structures to highlight discriminative patterns. Experiments on real‑world benchmarks show that SETTer outperforms state‑of‑the‑art models in 88% of scenarios.

By Abraham Ezema, Chijioke Eze, Ferdinanda Ponci, Antonello Monti
arXiv AI
Sep 10

Fewer yet critical: Reducing Redundant Token Dependencies for Transformer-based Time Series Forecasting

The paper introduces a token dependency selection strategy for Transformer-based time series forecasting. By jointly applying an attention entropy constraint and a prediction error constraint, the method identifies fewer but more critical inter-token dependencies, reducing the influence of redundant dependencies that can hurt generalization. Experiments on multiple datasets show that this approach improves forecasting performance across various Transformer models.

By Jianqi Zhang, Yuchan Liu, Zeen Song, Yuefei Li, Fanjiang Xu
Hugging Face Trending Papers
Sep 17

SETTer: Sparse-Encoder Transformer for Long-term Multivariate Time Series Forecasting

SETTer is a transformer-based model designed for long‑term multivariate time‑series forecasting. It introduces decoupled self‑attention and hybrid masking to better capture short‑ and long‑term patterns across time and channel dimensions, while adding simple explainable structures to highlight discriminative patterns. Experiments on real‑world benchmarks show that a single‑layer SETTer outperforms state‑of‑the‑art models in 88% of scenarios.

arXiv Machine Learning
2d ago

Aurora-X: Built for Extreme Time Series Forecasting

Aurora‑X is a billion‑parameter time‑series foundation model designed for extreme forecasting tasks. It employs a progressive curriculum that starts with channel‑independent pretraining, then adds cross‑variable dependencies, variable context and horizon lengths, and optional future covariates during mid‑training. A variable‑resolution post‑training stage allows adjustable temporal spans per token at inference, while a pattern‑guided mixture‑of‑experts expands capacity through sparse activation and expert specialization. An implicit quantile network head predicts arbitrary quantiles, enhancing probabilistic forecasting flexibility. Experiments on GIFT‑Eval, TIME, FEV‑Bench, TFB, and DAG‑Bench show state‑of‑the‑art performance against both pretrained TSFMs and task‑specific supervised models.

By Xingjian Wu, Chenjuan Guo, Xiangfei Qiu, Zhigang Hu, Hanyin Cheng, Peng Chen, Yang Shu, Jilin Hu, Bin Yang
arXiv AI
1d ago

Channel-Dependent State Space Model for Multivariate Time Series Forecasting

The paper introduces Chameleon, a channel‑dependent state space model for multivariate time series forecasting that allows data‑dependent, fine‑grained interactions across variables while maintaining linear scaling with the number of variables. By integrating selective state space models with a Kalman filter and adapting GatedDeltaNet as the backbone, Chameleon improves generalization and achieves lower MSE and MAE on strongly dependent ODE and PEMS datasets compared to both channel‑independent and prior channel‑dependent methods. Across 28 benchmark settings, it outperforms baselines in the majority of cases and demonstrates competitive training‑time and memory efficiency on Traffic and ETT datasets.

By Yu-Cheng Wu, Fan-Keng Sun, Li-Chun Lu, Duane S. Boning