Mixture-of-Experts (MoE) scales language models by routing each input through a small set of independently parameterized experts. We show that copying this design into convolutional networks fails for...
OMP-MoE is a training‑free compression framework that prunes redundant experts in Mixture‑of‑Experts large language models by framing the problem as sparse signal reconstruction solved with Orthogonal Matching Pursuit. The method greedily selects expert contributions as dictionary atoms to minimize reconstruction error, then optimizes cross‑layer expert allocation via a water‑filling strategy, and finally introduces an adaptive inference mechanism (OMP‑MoE†) that dynamically adjusts expert activation based on energy prediction. Experiments on Qwen, DeepSeek‑V2, GPT‑OSS, and Mixtral MoE show consistent performance gains at 25‑50% pruning ratios, with Qwen3‑30B‑A3B retaining 93.3% of original performance at 50% compression while achieving significant speedups.
By Dezhi Li, Lujun Li, Qiyuan Zhu, Hao Gu, Bei Liu, Sirui Han, Yike Guo
arXiv:2609.25655v1 Announce Type: new
Abstract: As large language models (LLMs) scale rapidly, dense full-parameter adaptation becomes increasingly expensive, motivating sparse and modular architectu...
By Zhentao Tan, Chang Liu, Yao Liu, Yue Wu, Jieping Ye
arXiv:2607. 29462v1 Announce Type: cross Abstract: Adapting deep learning models to profound clinical heterogeneity typically relies on parameter-efficient fine-tuning (PEFT) to avoid the severe overfitting associated with full end-to-end network updates.
By Sebastian Doerrich, Daniel W\"urtinger, Francesco Di Salvo, Shyam Nandan Rai, Christian Ledig
arXiv:2606. 27866v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models.
By Fan Mo, Yuxuan Han, Geng Zhang, Wangbo Zhao, Yang You
As large language models (LLMs) scale rapidly, dense full-parameter adaptation becomes increasingly expensive, motivating sparse and modular architectures such as Mixture-of-Experts (MoE) models. This...
arXiv:2609.09241v1 Announce Type: cross
Abstract: Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large f...
By Dohyeon Kim, Bedionita Soro, Sung Ju Hwang
IntBMoE introduces a block‑conditioned mixture‑of‑experts that decouples participation, execution, and materialization by combining dense expert composition with sparse block execution. Each internal layer uses a lightweight hypernetwork to merge all expert bases into a single composed expert, while a router selects only a few blocks per token, keeping compute and memory costs low. Experiments on image classification, language modeling, and sequential recommendation demonstrate consistent performance gains, and the model is deployed in AMap’s generative recommendation system, improving UVCTR by 2.4% in online A/B tests.
By Ran Cheng, Longfei Xu, Zheng Liu, Kaikui Liu, Xiangxiang Chu
MetaNet is a support‑set controller that predicts, for each layer of a Mixture‑of‑Experts model, an expert‑retention threshold and a bounded routing bias while keeping the backbone, experts, and router frozen. On DeepSeek‑MoE‑16B‑Chat, MetaNet offers a tunable trade‑off between accuracy and expert activation: a conservative setting activates 3.61 experts on average (40% fewer than a fixed k=6) with comparable MMLU accuracy, whereas an aggressive setting activates only 2.28 experts (62% fewer) with a modest accuracy drop. The MMLU‑trained controller also transfers to C‑Eval, activating 2.90 experts on average (52% fewer than fixed k=6) at 0.386 accuracy.
By Rongfeng Wang, Shichao Weng, Zhiqiang Wang, Xinyu Liu, Yang Yi, Peilong Zhou, Hongwei Tang
The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.
By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
arXiv:2607. 22577v1 Announce Type: new Abstract: Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training and inference costs grow linearly with model size-a critical bottleneck as models approach trillion-parameter regimes.
By Xin Yang, Yemin Wang, Mingda Liu, Letian Li, Shuaishuai Cao, Zhengxiao He, Ryan Dong
arXiv:2602. 06154v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) models scale large language models efficiently by sparsely activating experts, but once an expert is selected, it is executed fully.
By Nurbek Tastan, Stefanos Laskaridis, Karthik Nandakumar, Samuel Horvath