arXiv AI

Distributionally Robust Mixture-of-Experts Training

arXiv Machine Learning
3d ago

Dense Mixture-of-Experts as a Reparameterized Wide FFN: A Granularity Sweep at Fixed Compute

The paper investigates a dense analogue of Sparse Mixture-of-Experts (MoE) models by using $K$ SwiGLU experts that are all active for every token and combined via a softmax gate, keeping the total feed‑forward network (FFN) width fixed. Validation loss shows a non‑monotonic relationship with $K$: $K=2$ slightly improves performance over the single‑expert baseline, while $K=4$ and $K=6$ degrade it. The study also finds that the gating mechanism remains largely soft and balanced, except for a near one‑hot routing in the first layer of the $K=4$ model, which is functionally important as forcing uniform routing increases loss significantly.

By Vu Quang Hoang, Nghia Hieu Nguyen
arXiv Computation and Language
Aug 28

Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs

MetaNet is a support‑set controller that predicts, for each layer of a Mixture‑of‑Experts model, an expert‑retention threshold and a bounded routing bias while keeping the backbone, experts, and router frozen. On DeepSeek‑MoE‑16B‑Chat, MetaNet offers a tunable trade‑off between accuracy and expert activation: a conservative setting activates 3.61 experts on average (40% fewer than a fixed k=6) with comparable MMLU accuracy, whereas an aggressive setting activates only 2.28 experts (62% fewer) with a modest accuracy drop. The MMLU‑trained controller also transfers to C‑Eval, activating 2.90 experts on average (52% fewer than fixed k=6) at 0.386 accuracy.

By Rongfeng Wang, Shichao Weng, Zhiqiang Wang, Xinyu Liu, Yang Yi, Peilong Zhou, Hongwei Tang
arXiv AI
Sep 7

When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models

The paper examines how expert pruning—removing low‑importance experts in Mixture‑of‑Experts models—fails when the router is over‑dispersed, a condition caused by aggressive load‑balancing that spreads tokens nearly uniformly across experts. In this regime, traditional importance signals from router probabilities collapse, making perplexity an unreliable predictor of downstream accuracy; for example, the lowest‑perplexity pruning on gpt‑oss‑20B harms mathematical reasoning while the highest‑perplexity pruning preserves it. To address this, the authors introduce Minimax Expert Score Allocation (MESA), a domain‑aware method that iteratively boosts scores for the most affected domain, achieving minimal worst‑case degradation across domains and outperforming baseline pruning strategies on multiple benchmarks while reducing memory usage.

By Berkcan Kapusuzoglu, Connor Pryor, Sangwoo Cho, Supriyo Chakraborty, Shi-Xiong Zhang, Sambit Sahu, Milind Naphade
arXiv Machine Learning
6d ago

ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control

The paper introduces ID Balancing, an Integral‑Derivative controller that improves expert load balance in extremely sparse Mixture‑of‑Experts (MoE) models. By scaling its integral term with load error and activating the derivative term only when imbalance worsens, ID Balancing achieves over 50% reduction in worst‑case backbone MaxVio and 12% reduction in training‑average backbone MinVio compared to leading baselines, while preserving language‑modeling performance across various routing settings. The method’s benefits grow with increased sparsity, making it a promising approach for scaling larger MoE models.

By Peng Jin, Zihan Qiu, Zekun Wang, Bo Zheng, Yang Xu, Tian Xie, Xiao Li, Huaqing Zhang, Haoran Lian, Rui Men, Dayiheng Liu
arXiv Machine Learning
1d ago

MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling

MASKerade is a new dense‑to‑MoE training method that learns experts as sparse subnetworks of a frozen pretrained feed‑forward network (FFN) using binary masks. A token‑level router selects which masked FFNs to execute, and both the router and mask scores are jointly optimized while the underlying FFN weights remain unchanged. In experiments on five vision‑language benchmarks with Qwen and Gemma backbones, a configuration of four 2:4 experts with top‑2 routing achieved the best performance among compared baselines, demonstrating that mask learning over frozen weights is a practical alternative for constructing token‑routed MoE experts.

By Mingyuan Zhang, Yue Bai, Zhongruo Wang, Yupin Huang, Yiyang Huang, Hailing Wang, Huimin Zeng, Yun Fu