Federated adaptation of time-series foundation models (TSFMs) is attractive for building energy forecasting because meter data are private, distributed, and highly non-IID. However, a single parameter-sharing strategy is unlikely to serve all pretrained TSFMs or building clients: fully shared adapters can suppress building-specific temporal behavior, while fully local adaptation discards cross-building transfer.
OceanMoE is a structured conditional sparse Mixture-of-Experts framework designed for long‑horizon multivariate ocean forecasting. It fuses cross‑variable information to build target‑specific local representations and performs content‑conditioned sparse routing at each spatial location, with the number of active experts adjusted by router confidence. Experiments on ORAS5 data show that OceanMoE reduces aggregate forecasting error and maintains lower geometric‑mean normalized RMSE compared to baselines, while expert allocation varies with prediction targets and locations.
By Yishun Zhu, Jian Wang
arXiv:2606. 11625v1 Announce Type: new Abstract: Time-series foundation models (TSFMs) are increasingly explored as predictive experts within emerging agentic time-series systems.
By Kanghui Ning, Yushan Jiang, Kashif Rasul, Anderson Schneider, Yuriy Nevmyvaka, Dongjin Song
arXiv:2607. 06607v1 Announce Type: cross Abstract: Accurate long-term forecasting in complex systems is frequently compromised by dataset-level distribution shifts, where diverse underlying behavioral modes and evolving system states drive the dynamic multivariate time-series.
By Lanhao Li, Bingshu Xie, Lijun Sun, Xin Xue, Haoyi Zhou, Jianxin Li
arXiv:2601. 16632v4 Announce Type: replace-cross Abstract: Time series forecasting has witnessed significant progress with deep learning.
By Haonan Yang, Jianchao Tang, Zhuo Li
Aurora‑X is a billion‑parameter time‑series foundation model designed for extreme forecasting tasks. It employs a progressive curriculum that starts with channel‑independent pretraining, then adds cross‑variable dependencies, variable context and horizon lengths, and optional future covariates during mid‑training. A variable‑resolution post‑training stage allows adjustable temporal spans per token at inference, while a pattern‑guided mixture‑of‑experts expands capacity through sparse activation and expert specialization. An implicit quantile network head predicts arbitrary quantiles, enhancing probabilistic forecasting flexibility. Experiments on GIFT‑Eval, TIME, FEV‑Bench, TFB, and DAG‑Bench show state‑of‑the‑art performance against both pretrained TSFMs and task‑specific supervised models.
By Xingjian Wu, Chenjuan Guo, Xiangfei Qiu, Zhigang Hu, Hanyin Cheng, Peng Chen, Yang Shu, Jilin Hu, Bin Yang
arXiv:2607. 16882v1 Announce Type: new Abstract: Time series forecasting (TSF) is vital to many applications, yet existing models often struggle to capture the heterogeneous long-range global patterns and short-range local variations in multivariate time series.
By Wenqiang Ma, Chen Cheng, Xue Cheng, Jiarui Ye
arXiv:2607. 26618v1 Announce Type: new Abstract: Federated PEFT enables LLMs to collaboratively adapt to decentralized private data without sharing raw examples.
By Donghang Duan, Xu Zheng, Lizong Zhang, Chong Mu, Meng Han
The paper introduces DUMoE, a drift‑aware multimodal user representation framework that models user preferences over time by integrating static profiles, short‑term signals, and long‑term dependencies. It employs a sparse mixture‑of‑experts interest adapter, where each expert captures a distinct latent interest and a gating network selects relevant experts for each user. A three‑stage training strategy decouples backbone learning, expert specialization, and gating optimization, and experiments on real social media data demonstrate that DUMoE outperforms existing methods in user interest and interaction prediction.
By Ziqing Qian, Haohang Chen, Shengqi Dang, Yuhan Xiong, Canyu Shen, Jiaying Lei, Nan Cao
arXiv:2607. 09537v1 Announce Type: new Abstract: Time series forecasting requires models to capture diverse, often mutually exclusive, temporal dynamics, from smooth trend continuation to nonstationary drift and strict phase-aligned recurrence.
By Qitai Tan, Ruiwen Gu, Yilin Su, Mo Li, Xu Lin, Xiao-Ping Zhang
The paper introduces an adaptive Mixture-of-Experts (MoE) framework for time series forecasting that incorporates expert-specific losses to give each expert a direct learning signal independent of gating weights. The overall objective combines base forecasting loss with these expert losses, encouraging experts to specialize on different temporal segments. A partial online learning strategy is added for efficient incremental updates, and experiments on economic, tourism, and energy datasets show the method outperforms state‑of‑the‑art neural models and foundation models, with ablation studies confirming the benefit of expert loss integration.
By Btissame El Mahtout, Florian Ziel
Federated PEFT enables LLMs to collaboratively adapt to decentralized private data without sharing raw examples. However, task heterogeneity across clients can cause cross-task interference and gradient conflicts during aggregation.