arXiv Machine Learning By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang

Towards a Statistical Understanding of Mixture-of-Experts

Read the original on arXiv Machine Learning →

The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 17

MoRE: Mixture of Reused Experts

MoRE: Mixture of Reused Experts is a hybrid architecture that combines Mixture-of-Experts (MoE) with weight‑sharing techniques. It shares expert pools across adjacent layers while each layer keeps its own router, and introduces lightweight depth embeddings to help shared experts differentiate layer contexts. Experiments on models ranging from 114 M to 1.15 B parameters show MoRE achieves lower perplexity and better downstream performance than standard MoEs and other weight‑sharing models, with only minimal changes to existing MoE implementations.

By Eric S. Qiu, Utku Umur Acikalin, Justin Lovelace, Christian Belardi, Arjun B. Mulchandani, Carla P. Gomes, Kilian Q. Weinberger
arXiv AI
Jun 3

DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training

arXiv:2512. 13996v3 Announce Type: replace Abstract: Sparse Mixture-of-Experts architectures are essential for scaling model capacity efficiently, yet the standard Top-$k$ routing imposes a rigid sparsity pattern that ignores the intrinsic variance in token difficulty and layer-specific computational needs.

By Can Jin, Hongwu Peng, Mingcan Xiang, Qixin Zhang, Xiangchi Yuan, Amit Hasan, Ohi Dibua, Yifan Gong, Yan Kang, Dimitris N. Metaxas
arXiv AI
Sep 3

Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-Experts

The paper investigates how routing decisions in sparse mixture-of-experts (MoE) models evolve across layers. By aligning router control subspaces with generalized orthogonal Procrustes analysis, the authors find that a single linear transition can predict routing states across depth with substantial accuracy, revealing a shared geometric structure. They further demonstrate that these canonical states preserve expert selection better than generic hidden representations and improve next‑step routing predictions, reducing negative log‑likelihood by up to 15.7% on OLMoE and 6.2% on Phi.

By Kirill Labzin, Stepan Kulibaba, Artem Dzhalilov, Artem Gorokhov