The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.
By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
arXiv:2609. 17940v1 Announce Type: new Abstract: Sparse mixture-of-experts models route each token through a sequence of expert selections.
By Hao Li, Yasuyuki Tahara, Yuichi Sei
arXiv:2607. 28308v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions.
By Huiyuan Tian, Bonan Xu, Shijian Li
Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction.
arXiv:2608. 10392v1 Announce Type: new Abstract: Mixture-of-experts (MoE) models have recently moved beyond routing a fixed number of complete experts.
By Gongli Zhang, Zhulin Liu, C. L. Philip Chen
arXiv:2609.08189v1 Announce Type: new
Abstract: Dynamic layer routing reduces the inference cost of Large Language Models (LLMs) by learning to skip layers for individual tokens. Existing methods, ho...
By Hongjin Lin, Wentao Wan, Keze Wang
The paper introduces GRIP, an algorithm‑agnostic framework for machine unlearning in Mixture‑of‑Experts large language models. GRIP enforces hard geometric constraints on router updates, projecting gradient changes into the null space of the retain set’s routing matrix to prevent routing manipulation. Two variants—training‑time stochastic projection and post‑training analytical correction—show significant improvements in routing stability, retain accuracy, and resistance to white‑box adversarial recovery across two MoE models.
By Andy Zhu, Rongzhe Wei, Yupu Gu, Pan Li
arXiv:2512. 13996v3 Announce Type: replace Abstract: Sparse Mixture-of-Experts architectures are essential for scaling model capacity efficiently, yet the standard Top-$k$ routing imposes a rigid sparsity pattern that ignores the intrinsic variance in token difficulty and layer-specific computational needs.
By Can Jin, Hongwu Peng, Mingcan Xiang, Qixin Zhang, Xiangchi Yuan, Amit Hasan, Ohi Dibua, Yifan Gong, Yan Kang, Dimitris N. Metaxas
arXiv:2609.36301v1 Announce Type: cross
Abstract: Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regim...
By Honam Wong, Surbhi Goel, Enric Boix-Adser\`a
arXiv:2609.13058v1 Announce Type: new
Abstract: Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models hav...
By Hongyi He, Zhenghao Lin, Xiao Liu, Peng Cheng, Yan Lu, Yeyun Gong
arXiv:2604. 00421v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a learned router to map hidden states to expert assignments.
By Jama Hussein Mohamud, Drew Wagner, Mirco Ravanelli
arXiv:2606. 08814v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) scales model capacity efficiently by selectively routing inputs to a specialized subset of experts.
By Sumin Park, Noseong Park