arXiv Machine Learning

Fast Training of Mixture-of-Experts for Time Series Forecasting via Expert Loss Integration

The paper introduces an adaptive Mixture-of-Experts (MoE) framework for time series forecasting that incorporates expert-specific losses to give each expert a direct learning signal independent of gating weights. The overall objective combines base forecasting loss with these expert losses, encouraging experts to specialize on different temporal segments. A partial online learning strategy is added for efficient incremental updates, and experiments on economic, tourism, and energy datasets show the method outperforms state‑of‑the‑art neural models and foundation models, with ablation studies confirming the benefit of expert loss integration.

arXiv Machine Learning
Aug 31

Generalized Gibbs Ensemble Weighting for Forecast Combination

The paper introduces Generalized Gibbs Ensemble Weighting (GGEW), a probabilistic framework that assigns weights to forecasting models using a Gibbs-style exponential transformation of normalized predictive loss. GGEW extends basic weighting through numerical stabilization, diversity-aware score corrections, and online hyperparameter adaptation, yielding variants such as Stable Gibbs weighting, Directional Gibbs-NCL, and Symmetric Gibbs-NCL. The authors evaluate GGEW on M4 competition submissions and real-world datasets (Monash Traffic, Electricity, Solar), finding that Gibbs-style adaptive weighting is competitive across various settings, though performance varies by dataset, horizon, and deployment protocol.

By Prasen R. Nuthanakaluva, Nava K. Gaddam
arXiv Machine Learning
Sep 10

Electricity Price Forecasting: Bridging Linear Models, Neural Networks and Online Learning

The paper presents a hybrid neural architecture that blends linear and nonlinear feed‑forward networks for day‑ahead electricity price forecasting. It introduces a partial online learning strategy with warm‑starting and stage‑specific hyperparameters to cut computational time, and employs Bernstein Online Aggregation to combine forecasts. Experiments on six years of major European markets show the method reduces RMSE by 11‑12% and MAE by 14‑17% compared to state‑of‑the‑art benchmarks while lowering computational cost.

By Btissame El Mahtout, Florian Ziel
arXiv AI
Sep 7

Multi-Modal Time Series Prediction via Mixture of Modulated Experts

The paper introduces Expert Modulation, a novel approach for multi‑modal time series prediction that conditions both expert routing and computation on textual signals, thereby providing direct cross‑modal control over expert behavior. Unlike previous methods that rely on token‑level fusion, this mechanism avoids mixing temporal patches with language tokens in a shared embedding space, which can be problematic when high‑quality time‑text pairs are scarce or when time series characteristics vary widely. Experiments and theoretical analysis demonstrate that Expert Modulation yields strong improvements over existing multi‑modal forecasting techniques.

By Lige Zhang, Ali Maatouk, Jialin Chen, Karthik Charan Konduri, Leandros Tassiulas, Rex Ying
Hugging Face Trending Papers
Aug 19

An Empirical Benchmark of Deep Time-Series Models for Smart Meter Energy Forecasting

The paper presents an empirical benchmark of nine deep learning models for smart meter energy forecasting, evaluating them on two public datasets. It examines how historical input length, prediction horizon, and model architecture affect accuracy, finding that longer historical context improves performance up to a saturation point while accuracy declines with longer horizons. The study also compares computational cost, showing lightweight models achieve similar accuracy to heavier ones, and notes that model choice matters less across most population segments.

arXiv AI
Jun 3

AdaWeather: Adaptively Mixing Probabilistic Weather Forecasts with Logarithmic Regret

arXiv:2606. 02663v1 Announce Type: cross Abstract: Recent advances in machine learning have produced probabilistic weather forecasting models comparable to state-of-the-art numerical weather predictors.

By Saptarishi Dhanuka (Ashoka University), Sarvesh Iyer (Ashoka University), Manmeet Singh (Western Kentucky University), Mihir More (Ashoka University), Rushil Gupta (Ashoka University), Dhruman Gupta (Ashoka University), Parthasarathi Mukhopadhyay (Ashoka University), Sandeep Juneja (Ashoka University)
Hugging Face Trending Papers
Jun 8

FAME: Forecastability-Aware Mixture of Experts for Heterogeneous Time Series Forecasting

Large-scale retail and industrial forecasting systems contain many heterogeneous time series whose lifecycle, sparsity, volatility, seasonality, spectral patterns, and contextual sensitivity differ substantially. A single forecasting model rarely performs well across all regimes, while dense ensembles increase inference cost and provide limited insight into expert suitability.