Optimizing accuracy and diversity: a multi-task approach to forecast combinations
arXiv:2310. 20545v3 Announce Type: replace Abstract: We present a multi-task optimization approach based on a deep learning architecture for time series forecasting.
The paper introduces an adaptive Mixture-of-Experts (MoE) framework for time series forecasting that incorporates expert-specific losses to give each expert a direct learning signal independent of gating weights. The overall objective combines base forecasting loss with these expert losses, encouraging experts to specialize on different temporal segments. A partial online learning strategy is added for efficient incremental updates, and experiments on economic, tourism, and energy datasets show the method outperforms state‑of‑the‑art neural models and foundation models, with ablation studies confirming the benefit of expert loss integration.
arXiv:2310. 20545v3 Announce Type: replace Abstract: We present a multi-task optimization approach based on a deep learning architecture for time series forecasting.
arXiv:2510. 16898v2 Announce Type: replace-cross Abstract: Accurate prediction of electricity prices is crucial for stakeholders in the energy market, particularly for grid operators, energy producers, and consumers.
The paper introduces Generalized Gibbs Ensemble Weighting (GGEW), a probabilistic framework that assigns weights to forecasting models using a Gibbs-style exponential transformation of normalized predictive loss. GGEW extends basic weighting through numerical stabilization, diversity-aware score corrections, and online hyperparameter adaptation, yielding variants such as Stable Gibbs weighting, Directional Gibbs-NCL, and Symmetric Gibbs-NCL. The authors evaluate GGEW on M4 competition submissions and real-world datasets (Monash Traffic, Electricity, Solar), finding that Gibbs-style adaptive weighting is competitive across various settings, though performance varies by dataset, horizon, and deployment protocol.
The paper presents a hybrid neural architecture that blends linear and nonlinear feed‑forward networks for day‑ahead electricity price forecasting. It introduces a partial online learning strategy with warm‑starting and stage‑specific hyperparameters to cut computational time, and employs Bernstein Online Aggregation to combine forecasts. Experiments on six years of major European markets show the method reduces RMSE by 11‑12% and MAE by 14‑17% compared to state‑of‑the‑art benchmarks while lowering computational cost.
arXiv:2608. 12251v1 Announce Type: cross Abstract: Financial volatility is regime dependent, yet incorporating regime information into neural networks can also destabilize training.
arXiv:2606. 08896v1 Announce Type: new Abstract: Large-scale retail and industrial forecasting systems contain many heterogeneous time series whose lifecycle, sparsity, volatility, seasonality, spectral patterns, and contextual sensitivity differ substantially.
The paper introduces Expert Modulation, a novel approach for multi‑modal time series prediction that conditions both expert routing and computation on textual signals, thereby providing direct cross‑modal control over expert behavior. Unlike previous methods that rely on token‑level fusion, this mechanism avoids mixing temporal patches with language tokens in a shared embedding space, which can be problematic when high‑quality time‑text pairs are scarce or when time series characteristics vary widely. Experiments and theoretical analysis demonstrate that Expert Modulation yields strong improvements over existing multi‑modal forecasting techniques.
arXiv:2604. 22328v2 Announce Type: replace-cross Abstract: Driven by the transition towards a climate-neutral energy system, accurate energy time series forecasting is critical for planning and operations.
The paper presents an empirical benchmark of nine deep learning models for smart meter energy forecasting, evaluating them on two public datasets. It examines how historical input length, prediction horizon, and model architecture affect accuracy, finding that longer historical context improves performance up to a saturation point while accuracy declines with longer horizons. The study also compares computational cost, showing lightweight models achieve similar accuracy to heavier ones, and notes that model choice matters less across most population segments.
arXiv:2606. 02663v1 Announce Type: cross Abstract: Recent advances in machine learning have produced probabilistic weather forecasting models comparable to state-of-the-art numerical weather predictors.
arXiv:2608. 04695v1 Announce Type: cross Abstract: Federated adaptation of time-series foundation models (TSFMs) is attractive for building energy forecasting because meter data are private, distributed, and highly non-IID.
Large-scale retail and industrial forecasting systems contain many heterogeneous time series whose lifecycle, sparsity, volatility, seasonality, spectral patterns, and contextual sensitivity differ substantially. A single forecasting model rarely performs well across all regimes, while dense ensembles increase inference cost and provide limited insight into expert suitability.