arXiv AI

DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training

arXiv:2512. 13996v3 Announce Type: replace Abstract: Sparse Mixture-of-Experts architectures are essential for scaling model capacity efficiently, yet the standard Top-$k$ routing imposes a rigid sparsity pattern that ignores the intrinsic variance in token difficulty and layer-specific computational needs.

arXiv AI
4d ago

Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

The paper introduces two token-error supervision methods for sparse mixture-of-experts (MoE) large language models, aligning routing affinities with token-level cross-entropy loss. The first method predicts an error score per expert and uses it to adjust affinities before top‑K selection, while the second directly aligns router affinities to the model’s objective without extra heads or inference changes. Experiments on Granite and ARC‑Challenge show accuracy gains of about 2.3–2.94 percentage points over a parameter‑matched baseline, preserving the native sparse execution budget.

By Yury Nahshan, Nati Daniel, Jacob Goldberger, Yoli Shavit
arXiv Machine Learning
Sep 4

Towards a Statistical Understanding of Mixture-of-Experts

The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.

By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
arXiv Machine Learning
Aug 27

Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing

The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.

By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng
arXiv AI
Aug 5

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

arXiv:2608. 03457v1 Announce Type: new Abstract: Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood.

By Fengqi Zhu, Shaoxuan Xu, Jingyang Ou, Zebin You, Yipeng Xing, Huabin Liu, Xiaolu Zhang, Jun Zhou, Zhenzhong Lan, Yankai Lin, Wayne Xin Zhao, Jianguo Li, Chongxuan Li, Ji-Rong Wen