arXiv Machine Learning

Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers

The paper investigates how Mixture-of-Experts (MoE) Particle Transformers perform on the 188-class JetClass-II jet classification task. By varying expert count, routing capacity, top‑K, and auxiliary loss, the authors find that top‑1 MoE models can surpass dense baselines with similar nominal compute, but adding more experts yields diminishing accuracy gains. Activating multiple experts per token improves predictions at higher computational cost, and routing analyses reveal that expert assignments correlate with particle identity and kinematics, though this correlation does not consistently predict performance.

arXiv Machine Learning
Sep 22

Efficient Mixture-of-Experts with Speculative Decoding via Expert Coactivation

The paper studies how the design of Mixture-of-Experts (MoE) routers affects inference speed when combined with Speculative Decoding (SD). It shows that routers promoting high expert coactivation reduce memory transfer costs and improve runtime. By integrating a global load‑balancing loss, shared experts, a consistency loss, and an autoregressive expert selection mechanism, the authors achieve a 21% throughput gain over baseline MoEs while preserving accuracy.

By Kumari Nishu, Han-Byul Kim, Santosh Chilkunda, Maxwell Horton, Arnav Kundu, Mohammad Samragh, Lauren Hannah, Mohammad Sekhavat, Nikhil Bhendawade, Manuel Ciosici, Iman Mirzadeh, Keivan Alizadeh Vahid, David Harrison, Irina Belousova, Mehrdad Farajtabar, Minsik Cho
arXiv Machine Learning
Jul 24

PreMoE: Proactive Inference for Efficient Mixture-of-Experts

arXiv:2505. 17639v4 Announce Type: replace Abstract: Mixture-of-Experts (MoE) models offer dynamic computation, but are typically deployed as static full-capacity models, missing opportunities for deployment-specific specialization.

By Zehua Pei, Ying Zhang, Hui-Ling Zhen, Tao Yuan, Xianzhi Yu, Zhenhua Dong, Sinno Jialin Pan, Mingxuan Yuan, Bei Yu
arXiv AI
Jun 3

DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training

arXiv:2512. 13996v3 Announce Type: replace Abstract: Sparse Mixture-of-Experts architectures are essential for scaling model capacity efficiently, yet the standard Top-$k$ routing imposes a rigid sparsity pattern that ignores the intrinsic variance in token difficulty and layer-specific computational needs.

By Can Jin, Hongwu Peng, Mingcan Xiang, Qixin Zhang, Xiangchi Yuan, Amit Hasan, Ohi Dibua, Yifan Gong, Yan Kang, Dimitris N. Metaxas
arXiv Machine Learning
Aug 27

Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing

The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.

By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng