arXiv Machine Learning By Yushuai Sun, Zikun Zhou, Lin Gao, Jun Yu, Wenjie Pei

MoRA: MoE Pruning via Router Bias Learning and Expert Approximation

Read the original on arXiv Machine Learning →

MoRA is a framework for pruning Mixture-of-Experts (MoE) models by learning a router bias for each expert and optimizing it with a language‑modeling loss and a routing‑diversity regularizer. The learned biases sharpen routing distributions to identify critical experts and encourage diverse routing preferences. After pruning, MoRA uses an expert approximation mechanism that approximates the outputs of pruned experts with affine transformations of remaining experts, further improving performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 27

Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing

The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.

By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng
arXiv AI
Sep 7

When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models

The paper examines how expert pruning—removing low‑importance experts in Mixture‑of‑Experts models—fails when the router is over‑dispersed, a condition caused by aggressive load‑balancing that spreads tokens nearly uniformly across experts. In this regime, traditional importance signals from router probabilities collapse, making perplexity an unreliable predictor of downstream accuracy; for example, the lowest‑perplexity pruning on gpt‑oss‑20B harms mathematical reasoning while the highest‑perplexity pruning preserves it. To address this, the authors introduce Minimax Expert Score Allocation (MESA), a domain‑aware method that iteratively boosts scores for the most affected domain, achieving minimal worst‑case degradation across domains and outperforming baseline pruning strategies on multiple benchmarks while reducing memory usage.

By Berkcan Kapusuzoglu, Connor Pryor, Sangwoo Cho, Supriyo Chakraborty, Shi-Xiong Zhang, Sambit Sahu, Milind Naphade
arXiv Machine Learning
4d ago

OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit

OMP-MoE is a training‑free compression framework that prunes redundant experts in Mixture‑of‑Experts large language models by framing the problem as sparse signal reconstruction solved with Orthogonal Matching Pursuit. The method greedily selects expert contributions as dictionary atoms to minimize reconstruction error, then optimizes cross‑layer expert allocation via a water‑filling strategy, and finally introduces an adaptive inference mechanism (OMP‑MoE†) that dynamically adjusts expert activation based on energy prediction. Experiments on Qwen, DeepSeek‑V2, GPT‑OSS, and Mixtral MoE show consistent performance gains at 25‑50% pruning ratios, with Qwen3‑30B‑A3B retaining 93.3% of original performance at 50% compression while achieving significant speedups.

By Dezhi Li, Lujun Li, Qiyuan Zhu, Hao Gu, Bei Liu, Sirui Han, Yike Guo
arXiv Machine Learning
Jun 16

How to Score Experts for One-Shot MoE Expert Pruning: A Unified Formulation and Selection Principle

arXiv:2606. 15716v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models reduce per-token computation through sparse expert activation, yet deployment still requires storing the full expert pool, making one-shot expert pruning a practical approach for reducing memory usage.

By Zongfang Liu, Jinghui Zhang, Zijian Ma, Guangyi Chen, Xin Yuan