arXiv AI

SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs

arXiv:2606. 17952v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (MoE) architectures enable scaling LLM parameters under a fixed inference budget by activating only a small subset of experts via top-$k$ routing.

arXiv Machine Learning
Sep 4

Towards a Statistical Understanding of Mixture-of-Experts

The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.

By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
arXiv Machine Learning
Aug 27

Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing

The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.

By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng
arXiv AI
Jun 3

DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training

arXiv:2512. 13996v3 Announce Type: replace Abstract: Sparse Mixture-of-Experts architectures are essential for scaling model capacity efficiently, yet the standard Top-$k$ routing imposes a rigid sparsity pattern that ignores the intrinsic variance in token difficulty and layer-specific computational needs.

By Can Jin, Hongwu Peng, Mingcan Xiang, Qixin Zhang, Xiangchi Yuan, Amit Hasan, Ohi Dibua, Yifan Gong, Yan Kang, Dimitris N. Metaxas
arXiv AI
Jun 2

DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts

arXiv:2606. 01062v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have become a leading approach for decoupling parameter count from computational cost in large language models, yet effectively scaling MoE performance remains a challenge.

By Jiarui Feng, Hanqing Zeng, Karish Grover, Ruizhong Qiu, Yinglong Xia, Qiang Zhang, Qifan Wang, Ren Chen, Dongqi Fu, Jiayi Liu, Zhoukai Zhao, Xiangjun Fan, Benyu Zhang, Yixin Chen
arXiv Machine Learning
Sep 21

IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

IntBMoE introduces a block‑conditioned mixture‑of‑experts that decouples participation, execution, and materialization by combining dense expert composition with sparse block execution. Each internal layer uses a lightweight hypernetwork to merge all expert bases into a single composed expert, while a router selects only a few blocks per token, keeping compute and memory costs low. Experiments on image classification, language modeling, and sequential recommendation demonstrate consistent performance gains, and the model is deployed in AMap’s generative recommendation system, improving UVCTR by 2.4% in online A/B tests.

By Ran Cheng, Longfei Xu, Zheng Liu, Kaikui Liu, Xiangxiang Chu
arXiv AI
Sep 17

MoRE: Mixture of Reused Experts

MoRE: Mixture of Reused Experts is a hybrid architecture that combines Mixture-of-Experts (MoE) with weight‑sharing techniques. It shares expert pools across adjacent layers while each layer keeps its own router, and introduces lightweight depth embeddings to help shared experts differentiate layer contexts. Experiments on models ranging from 114 M to 1.15 B parameters show MoRE achieves lower perplexity and better downstream performance than standard MoEs and other weight‑sharing models, with only minimal changes to existing MoE implementations.

By Eric S. Qiu, Utku Umur Acikalin, Justin Lovelace, Christian Belardi, Arjun B. Mulchandani, Carla P. Gomes, Kilian Q. Weinberger
arXiv Machine Learning
Jul 15

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs

arXiv:2602. 19938v2 Announce Type: replace Abstract: Sparse Mixture-of-Experts (SMoE) architectures are increasingly used to scale large language models efficiently, delivering strong accuracy under fixed compute budgets.

By Zijie Liu, Jie Peng, Jinhao Duan, Zirui Liu, Kaixiong Zhou, Mingfu Liang, Luke Simon, Xi Liu, Zhaozhuo Xu, Tianlong Chen