The paper investigates a dense analogue of Sparse Mixture-of-Experts (MoE) models by using $K$ SwiGLU experts that are all active for every token and combined via a softmax gate, keeping the total feed‑forward network (FFN) width fixed. Validation loss shows a non‑monotonic relationship with $K$: $K=2$ slightly improves performance over the single‑expert baseline, while $K=4$ and $K=6$ degrade it. The study also finds that the gating mechanism remains largely soft and balanced, except for a near one‑hot routing in the first layer of the $K=4$ model, which is functionally important as forcing uniform routing increases loss significantly.
By Vu Quang Hoang, Nghia Hieu Nguyen
MetaNet is a support‑set controller that predicts, for each layer of a Mixture‑of‑Experts model, an expert‑retention threshold and a bounded routing bias while keeping the backbone, experts, and router frozen. On DeepSeek‑MoE‑16B‑Chat, MetaNet offers a tunable trade‑off between accuracy and expert activation: a conservative setting activates 3.61 experts on average (40% fewer than a fixed k=6) with comparable MMLU accuracy, whereas an aggressive setting activates only 2.28 experts (62% fewer) with a modest accuracy drop. The MMLU‑trained controller also transfers to C‑Eval, activating 2.90 experts on average (52% fewer than fixed k=6) at 0.386 accuracy.
By Rongfeng Wang, Shichao Weng, Zhiqiang Wang, Xinyu Liu, Yang Yi, Peilong Zhou, Hongwei Tang
The paper examines how expert pruning—removing low‑importance experts in Mixture‑of‑Experts models—fails when the router is over‑dispersed, a condition caused by aggressive load‑balancing that spreads tokens nearly uniformly across experts. In this regime, traditional importance signals from router probabilities collapse, making perplexity an unreliable predictor of downstream accuracy; for example, the lowest‑perplexity pruning on gpt‑oss‑20B harms mathematical reasoning while the highest‑perplexity pruning preserves it. To address this, the authors introduce Minimax Expert Score Allocation (MESA), a domain‑aware method that iteratively boosts scores for the most affected domain, achieving minimal worst‑case degradation across domains and outperforming baseline pruning strategies on multiple benchmarks while reducing memory usage.
By Berkcan Kapusuzoglu, Connor Pryor, Sangwoo Cho, Supriyo Chakraborty, Shi-Xiong Zhang, Sambit Sahu, Milind Naphade
The paper introduces ID Balancing, an Integral‑Derivative controller that improves expert load balance in extremely sparse Mixture‑of‑Experts (MoE) models. By scaling its integral term with load error and activating the derivative term only when imbalance worsens, ID Balancing achieves over 50% reduction in worst‑case backbone MaxVio and 12% reduction in training‑average backbone MinVio compared to leading baselines, while preserving language‑modeling performance across various routing settings. The method’s benefits grow with increased sparsity, making it a promising approach for scaling larger MoE models.
By Peng Jin, Zihan Qiu, Zekun Wang, Bo Zheng, Yang Xu, Tian Xie, Xiao Li, Huaqing Zhang, Haoran Lian, Rui Men, Dayiheng Liu
arXiv:2607. 24665v1 Announce Type: cross Abstract: Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows.
By Yanhao Jia, Jiepeng Wang, Haibin Huang, Chi Zhang, Erik Cambria, Xuelong Li
arXiv:2609. 04575v1 Announce Type: cross Abstract: Modern fine-grained Mixture-of-Experts (MoE) models route each token to a small number of experts and renormalize their router probabilities.
By Xing Chen, Hengshuai Yao
arXiv:2606. 03391v1 Announce Type: cross Abstract: Model merging has emerged as a cost-effective approach for consolidating the capabilities of multiple LLMs without retraining.
By Canbin Huang, Tianyuan Shi, Xiaojun Quan, Jingang Wang, Jianfei Zhang, Qifan Wang
arXiv:2604.23036v2 Announce Type: replace-cross
Abstract: Despite MoE models leading many benchmarks, supervised fine-tuning (SFT) for the MoE architectures remains difficult because its router layer...
By Haoze He, Xingyuan Ding, Xuan Jiang, Xinkai Zou, Alex Cheng, Yibo Zhao, Juncheng Billy Li, Heather Miller
arXiv:2609.36222v1 Announce Type: new
Abstract: Large language models are increasingly expensive to serve. In large-scale serving systems, autoregressive decoding is often bottlenecked by transferrin...
By Ali Abbasi, Justin Shi, Soheil Kolouri
arXiv:2609.09241v1 Announce Type: cross
Abstract: Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large f...
By Dohyeon Kim, Bedionita Soro, Sung Ju Hwang
MASKerade is a new dense‑to‑MoE training method that learns experts as sparse subnetworks of a frozen pretrained feed‑forward network (FFN) using binary masks. A token‑level router selects which masked FFNs to execute, and both the router and mask scores are jointly optimized while the underlying FFN weights remain unchanged. In experiments on five vision‑language benchmarks with Qwen and Gemma backbones, a configuration of four 2:4 experts with top‑2 routing achieved the best performance among compared baselines, demonstrating that mask learning over frozen weights is a practical alternative for constructing token‑routed MoE experts.
By Mingyuan Zhang, Yue Bai, Zhongruo Wang, Yupin Huang, Yiyang Huang, Hailing Wang, Huimin Zeng, Yun Fu
arXiv:2609.36724v1 Announce Type: new
Abstract: Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gr...
By Yuchen Li, Mingyu Du, Zongqi Fan, Nguyen H. Tran, Ken-Tye Yong