arXiv Machine Learning

OceanMoE: Structured Conditional Sparse Computation for Long-Horizon Multivariate Ocean Forecasting

OceanMoE is a structured conditional sparse Mixture-of-Experts framework designed for long‑horizon multivariate ocean forecasting. It fuses cross‑variable information to build target‑specific local representations and performs content‑conditioned sparse routing at each spatial location, with the number of active experts adjusted by router confidence. Experiments on ORAS5 data show that OceanMoE reduces aggregate forecasting error and maintains lower geometric‑mean normalized RMSE compared to baselines, while expert allocation varies with prediction targets and locations.

arXiv Machine Learning
Sep 4

Towards a Statistical Understanding of Mixture-of-Experts

The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.

By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
Hugging Face Trending Papers
Aug 5

Personalized Federated Sparse Adaptation of Time-Series Foundation Models

Federated adaptation of time-series foundation models (TSFMs) is attractive for building energy forecasting because meter data are private, distributed, and highly non-IID. However, a single parameter-sharing strategy is unlikely to serve all pretrained TSFMs or building clients: fully shared adapters can suppress building-specific temporal behavior, while fully local adaptation discards cross-building transfer.

arXiv Machine Learning
Sep 11

M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction

M3-Former is a multimodal transformer framework that uses large language models to encode vessel static attributes and navigational intent as semantic priors for long‑term trajectory prediction. It builds a unified multimodal representation space, aligns static semantic information with dynamic trajectory features via self‑attention, and employs a dual‑granularity Mixture‑of‑Experts architecture to capture both global route planning and fine‑grained maneuvering behaviors. A Steering‑Weighted Cross‑Entropy loss further improves accuracy on sparse turning events, and experiments on a Danish AIS dataset show consistent improvements over state‑of‑the‑art baselines, reducing ADE and FDE by up to 5.1% in 4‑hour predictions.

By Wenzhe Jin, Haina Tang
arXiv Machine Learning
Jun 11

PCA-Enhanced Adaptive NVAR Framework for High-Resolution Sea Surface Temperature Forecasting in the East Sea

arXiv:2606. 12141v1 Announce Type: new Abstract: Accurate forecasting of sea surface temperature (SST) in regional seas such as the East Sea is crucial for monitoring marine ecosystems, assessing climate risks, managing fisheries, and conducting naval operations.

By Sherkhon Azimov, Susana L\'opez-Moreno, Eric Dolores-Cuenca, JinYong Choi, Sangil Kim
arXiv Machine Learning
Jul 21

BG4Sea: Biogeochemical Seasonal Forecastability via Progressive Information Scaling

arXiv:2607. 16731v1 Announce Type: new Abstract: Marine biogeochemical forecasting is increasingly important for managing marine ecosystems and the carbon cycle, yet global, seasonal forecast products lag far behind physical oceanography, held back by the complexity of the processes involved and by data scarcity.

By Gabriela Martinez Balbontin, Anastase Charantonis, Dominique Bereziat, Stefano Ciavatta
arXiv AI
Jul 22

Incomplete Observations Boost Evolutionary Performance in Ocean Modeling

arXiv:2607. 19147v1 Announce Type: cross Abstract: Data-driven methods have revolutionized ocean modeling, yet current approaches rely heavily on complete reanalysis datasets, imposing computational constraints and limiting model performance to that of the training data.

By Yangyang Kong, Yutong Jiang, Yanhai Gan, Junyu Dong, Feng Gao, Xiaopei Lin