arXiv Machine Learning

MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling

MASKerade is a new dense‑to‑MoE training method that learns experts as sparse subnetworks of a frozen pretrained feed‑forward network (FFN) using binary masks. A token‑level router selects which masked FFNs to execute, and both the router and mask scores are jointly optimized while the underlying FFN weights remain unchanged. In experiments on five vision‑language benchmarks with Qwen and Gemma backbones, a configuration of four 2:4 experts with top‑2 routing achieved the best performance among compared baselines, demonstrating that mask learning over frozen weights is a practical alternative for constructing token‑routed MoE experts.

arXiv Computation and Language
Aug 28

Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs

MetaNet is a support‑set controller that predicts, for each layer of a Mixture‑of‑Experts model, an expert‑retention threshold and a bounded routing bias while keeping the backbone, experts, and router frozen. On DeepSeek‑MoE‑16B‑Chat, MetaNet offers a tunable trade‑off between accuracy and expert activation: a conservative setting activates 3.61 experts on average (40% fewer than a fixed k=6) with comparable MMLU accuracy, whereas an aggressive setting activates only 2.28 experts (62% fewer) with a modest accuracy drop. The MMLU‑trained controller also transfers to C‑Eval, activating 2.90 experts on average (52% fewer than fixed k=6) at 0.386 accuracy.

By Rongfeng Wang, Shichao Weng, Zhiqiang Wang, Xinyu Liu, Yang Yi, Peilong Zhou, Hongwei Tang
arXiv AI
3d ago

MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference

MaskCoFT introduces a masked co‑adaptive fine‑tuning approach for mixture‑of‑experts language models, training both routers and experts jointly with a cross‑entropy loss. A learnable binary mask limits each layer’s Top‑K routing to a subset of experts during fine‑tuning, allowing experts to adapt to the tokens they receive. In simulated GPU cache scenarios, MaskCoFT reduces expert fetches per token by 23.7% for Mixtral‑8x7B and 10.1% for DeepSeek‑V2‑Lite, and lowers inference time per output token by up to 16.4% and 5.5% respectively, while maintaining or improving accuracy across nine benchmarks.

By Junfeng Wu, Zehao Fan, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu, Yinan Wang
arXiv Machine Learning
Aug 27

Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing

The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.

By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng
arXiv Machine Learning
Sep 4

Towards a Statistical Understanding of Mixture-of-Experts

The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.

By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
arXiv Machine Learning
Sep 21

IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

IntBMoE introduces a block‑conditioned mixture‑of‑experts that decouples participation, execution, and materialization by combining dense expert composition with sparse block execution. Each internal layer uses a lightweight hypernetwork to merge all expert bases into a single composed expert, while a router selects only a few blocks per token, keeping compute and memory costs low. Experiments on image classification, language modeling, and sequential recommendation demonstrate consistent performance gains, and the model is deployed in AMap’s generative recommendation system, improving UVCTR by 2.4% in online A/B tests.

By Ran Cheng, Longfei Xu, Zheng Liu, Kaikui Liu, Xiangxiang Chu
arXiv Machine Learning
3d ago

Dense Mixture-of-Experts as a Reparameterized Wide FFN: A Granularity Sweep at Fixed Compute

The paper investigates a dense analogue of Sparse Mixture-of-Experts (MoE) models by using $K$ SwiGLU experts that are all active for every token and combined via a softmax gate, keeping the total feed‑forward network (FFN) width fixed. Validation loss shows a non‑monotonic relationship with $K$: $K=2$ slightly improves performance over the single‑expert baseline, while $K=4$ and $K=6$ degrade it. The study also finds that the gating mechanism remains largely soft and balanced, except for a near one‑hot routing in the first layer of the $K=4$ model, which is functionally important as forcing uniform routing increases loss significantly.

By Vu Quang Hoang, Nghia Hieu Nguyen