arXiv:2606. 29347v1 Announce Type: cross Abstract: Adaptive Financial Transformer (AFT) is proposed for stock return prediction under non-stationary financial markets.
By Dishan Sarkar
arXiv:2602. 02288v3 Announce Type: replace Abstract: Current time-series forecasting models are primarily based on transformer-style neural networks.
By Zheng Li, Jerry Cheng, Huanying Gu
arXiv:2607. 15705v1 Announce Type: new Abstract: Accurate load forecasting at multiple grid levels is essential for future smart grids, ranging from aggregated control area forecasts for balancing supply and demand to forecasts of individual end-consumer loads for demand-side management and energy management systems.
By Matthias Hertel, Sebastian P\"utz, Jonathan Kolar, Benjamin Sch\"afer, Ralf Mikut, Veit Hagenmeyer
arXiv:2606. 10678v1 Announce Type: new Abstract: Transformer-based models have emerged as leading paradigms in time-series forecasting in recent years, employing self-attention mechanisms to capture long-range dependencies.
By Amrijit Biswas, Mustafa Kamal, Robin Krambroeckers, M. M. Lutfe Elahi, Sifat Momen, Nabeel Mohammed, Shafin Rahman
arXiv:2607. 00154v1 Announce Type: cross Abstract: Evolutionary neural architecture design for multivariate time-series forecasting remains underexplored, with most approaches relying on fixed Transformer architectures despite substantial variation across tasks and forecasting settings.
By AbdElRahman ElSaid, Damir Pulatov
The paper introduces an adaptive Mixture-of-Experts (MoE) framework for time series forecasting that incorporates expert-specific losses to give each expert a direct learning signal independent of gating weights. The overall objective combines base forecasting loss with these expert losses, encouraging experts to specialize on different temporal segments. A partial online learning strategy is added for efficient incremental updates, and experiments on economic, tourism, and energy datasets show the method outperforms state‑of‑the‑art neural models and foundation models, with ablation studies confirming the benefit of expert loss integration.
By Btissame El Mahtout, Florian Ziel
The paper introduces a token dependency selection strategy for Transformer-based time series forecasting. By jointly applying an attention entropy constraint and a prediction error constraint, the method identifies fewer but more critical inter-token dependencies, reducing the influence of redundant dependencies that can hurt generalization. Experiments on multiple datasets show that this approach improves forecasting performance across various Transformer models.
By Jianqi Zhang, Yuchan Liu, Zeen Song, Yuefei Li, Fanjiang Xu
arXiv:2505. 15354v3 Announce Type: replace Abstract: Time-series forecasting is a critical task in various business domains, but it remains inherently challenging.
By Hamza Cherkaoui, Malik Tiomoko, Giuseppe Paolo, Zhang Yili, Yu Meng, Zhang Keli, Hafiz Tiomoko Ali
The paper introduces a lightweight pre‑encoder gate for Transformer‑based time‑series forecasters, which assigns sigmoid scores to covariate representations before they enter the encoder. The gate is evaluated as a plug‑in for models such as TimeXer, iTransformer, and PatchTST on datasets including ETTm1, ETTm2, Traffic, Energy, and ILI, showing competitive performance and the ability to regulate covariate admission via a usage penalty. Experiments also explore gate placement, initialization, and feature importance using VIF‑informed permutation diagnostics.
By Hongkai Zhuang, Tao Huang, Chen Hou
arXiv:2511. 09789v3 Announce Type: replace Abstract: Short-term energy forecasting plays an important role in real-time operational decision-making, such as electricity market bidding and power system dispatch, where both numerical accuracy and correct directional signals are essential.
By Fulong Yao, Wanqing Zhao, Chao Zheng, Xiaofei Han
DualCast is a dual‑path language model that forecasts financial time‑series by combining a fast numerical forecaster with an optional text‑conditioned revision mechanism. The fast path trains only new financial‑token embeddings and output heads on a frozen Qwen3‑8B backbone, while the slow path uses a LoRA adapter to incorporate news and refine predictions. In zero‑shot tests across equities and energy prices at multiple time resolutions, the slow path achieves the lowest mean absolute percentage error in most settings, especially for longer horizons, and news ablations show additional gains in many markets.
By Wentao Zhao, Hongqiang Wu, Shanghang Liu, Zhaochen Zan, Yu Zhang, Biqing Huang
arXiv:2607. 22299v1 Announce Type: cross Abstract: Forecasting multiple time-series with high-dimensional covariates presents a core challenge: unifying common temporal patterns while retaining meaningful series-specific information.
By Wan Zhang, Qinjie Lin, Chan Lee, Weijian Li, Han Liu, Kai Zhang