arXiv Machine Learning

Regime-Gated Residual Mixture-of-Experts for Cross-Sectional Volatility Forecasting

arXiv:2608. 12251v1 Announce Type: cross Abstract: Financial volatility is regime dependent, yet incorporating regime information into neural networks can also destabilize training.

arXiv Machine Learning
Jul 13

Forking-Sequences: Statistically and Computationally Efficient Multi-Horizon Forecasting with Reduced Volatility

arXiv:2510. 04487v5 Announce Type: replace Abstract: While accuracy is a critical requirement for time series forecasting, an equally important desideratum is reasonable forecast volatility across forecast creation dates (FCDs).

By Willa Potosnak, Malcolm Wolff, Mengfei Cao, Ruijun Ma, Tatiana Konstantinova, Dmitry Efimov, Michael W. Mahoney, Boris Oreshkin, Kin G. Olivares
arXiv AI
Aug 19

MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting

MoFE is a deep learning framework that combines Fourier Neural Operators with a Mixture-of-Experts architecture to forecast cryptocurrency prices. It models volatility as a mix of multi-frequency components—including fundamental growth, mining costs, halving events, and market sentiment—using adaptive FNO and convolutional experts. Experiments on Bitcoin data from 2020 to 2025 show MoFE outperforms existing models in short‑term horizons, reducing phase‑lag errors and improving directional accuracy and information coefficient, which translates into higher Sharpe ratios in simulated trading.

By Bowen Liu, Mingming Sun
arXiv Machine Learning
Sep 18

Fast Training of Mixture-of-Experts for Time Series Forecasting via Expert Loss Integration

The paper introduces an adaptive Mixture-of-Experts (MoE) framework for time series forecasting that incorporates expert-specific losses to give each expert a direct learning signal independent of gating weights. The overall objective combines base forecasting loss with these expert losses, encouraging experts to specialize on different temporal segments. A partial online learning strategy is added for efficient incremental updates, and experiments on economic, tourism, and energy datasets show the method outperforms state‑of‑the‑art neural models and foundation models, with ablation studies confirming the benefit of expert loss integration.

By Btissame El Mahtout, Florian Ziel
arXiv Machine Learning
Sep 14

VertiFuseX: Generalizable Financial Forecasting via Multi-Stream Temporal Fusion

VertiFuseX is a hybrid LSTM architecture that fuses multi‑scale temporal representations at the penultimate layer, stacking features from LSTM, Bi‑LSTM, and St‑LSTM branches and a parallel DNN stream. On 15 years of global equity index data, it reduces MAPE by 30‑54% and improves MAE and RMSE by over 40% compared to LSTM baselines, outperforming seven state‑of‑the‑art models across 33 metric‑dataset comparisons. The model is lightweight (675k parameters, 2.6 MB footprint) with 1.5 ms/sample inference latency and demonstrates robust, interpretable forecasting with reduced drawdowns in algorithmic trading simulations.

By Aashish Bohra, Vivek Vijay
Hugging Face Trending Papers
Jun 8

FAME: Forecastability-Aware Mixture of Experts for Heterogeneous Time Series Forecasting

Large-scale retail and industrial forecasting systems contain many heterogeneous time series whose lifecycle, sparsity, volatility, seasonality, spectral patterns, and contextual sensitivity differ substantially. A single forecasting model rarely performs well across all regimes, while dense ensembles increase inference cost and provide limited insight into expert suitability.

arXiv Machine Learning
Aug 28

Graph-Based Modeling of Financial Volatility Dynamics

The paper introduces the Finance‑Aware Graph Spatio‑Temporal Network (FA‑GSTN) for forecasting realized volatility by treating the implied volatility surface as a dynamic graph. Nodes represent grid points on the surface, with edges capturing adaptive intra‑day spatial and explicit inter‑day temporal relationships, while finance‑aware node features (e.g., option Greeks) and a multi‑scale temporal smoothing gate address high‑frequency noise. Experiments on a large equity options dataset show FA‑GSTN achieves state‑of‑the‑art predictive accuracy (R² up to 0.473) and outperforms Vision Transformer baselines even with only one year of training data, demonstrating robustness during market stress.

By Chuanzhen Wang, Alice Zhang, Wei Chen, Michael Brown