arXiv:2607. 12696v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (MoE) models have become an important approach for scaling Large Language Models (LLMs), but their inference efficiency depends strongly on expert activation patterns.
By Jincheng Xie, Runheng Liu, Heyan Huang, Yawen Ling, Hanbin Dai, Yu Zheng, Wen Hu
arXiv:2609.08690v1 Announce Type: cross
Abstract: Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes reliable hyp...
By Changxin Tian, Kunlong Chen, Jia Liu, Ziqi Liu, Zhiqiang Zhang, Jun Zhou
ITC-MoE introduces an Importance-guided Token-aware Compression framework for Mixture-of-Experts Diffusion Language Models. It combines Adaptive Tucker Compression, which uses activation and gradient importance to jointly factorize expert weights and allocate ranks, with Token-aware Compensation and Routing that applies low‑rank adjustments to hot tokens and limits expert candidates for cold tokens. The method achieves significant reductions in computation and storage while maintaining generation quality, exemplified by a 30% compression budget that preserves 96.33% accuracy on MultiArith and delivers up to a 7.22× speedup.
By Lianjun Liu, Shipeng Li, You Huang, Weiqi Yan, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong
arXiv:2609.36222v1 Announce Type: new
Abstract: Large language models are increasingly expensive to serve. In large-scale serving systems, autoregressive decoding is often bottlenecked by transferrin...
By Ali Abbasi, Justin Shi, Soheil Kolouri
The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.
By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng
arXiv:2512. 13996v3 Announce Type: replace Abstract: Sparse Mixture-of-Experts architectures are essential for scaling model capacity efficiently, yet the standard Top-$k$ routing imposes a rigid sparsity pattern that ignores the intrinsic variance in token difficulty and layer-specific computational needs.
By Can Jin, Hongwu Peng, Mingcan Xiang, Qixin Zhang, Xiangchi Yuan, Amit Hasan, Ohi Dibua, Yifan Gong, Yan Kang, Dimitris N. Metaxas