arXiv Machine Learning By Mingyuan Zhang, Yue Bai, Zhongruo Wang, Yupin Huang, Yiyang Huang, Hailing Wang, Huimin Zeng, Yun Fu

MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling

Read the original on arXiv Machine Learning →

MASKerade is a new dense‑to‑MoE training method that learns experts as sparse subnetworks of a frozen pretrained feed‑forward network (FFN) using binary masks. A token‑level router selects which masked FFNs to execute, and both the router and mask scores are jointly optimized while the underlying FFN weights remain unchanged. In experiments on five vision‑language benchmarks with Qwen and Gemma backbones, a configuration of four 2:4 experts with top‑2 routing achieved the best performance among compared baselines, demonstrating that mask learning over frozen weights is a practical alternative for constructing token‑routed MoE experts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 28

Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs

MetaNet is a support‑set controller that predicts, for each layer of a Mixture‑of‑Experts model, an expert‑retention threshold and a bounded routing bias while keeping the backbone, experts, and router frozen. On DeepSeek‑MoE‑16B‑Chat, MetaNet offers a tunable trade‑off between accuracy and expert activation: a conservative setting activates 3.61 experts on average (40% fewer than a fixed k=6) with comparable MMLU accuracy, whereas an aggressive setting activates only 2.28 experts (62% fewer) with a modest accuracy drop. The MMLU‑trained controller also transfers to C‑Eval, activating 2.90 experts on average (52% fewer than fixed k=6) at 0.386 accuracy.

By Rongfeng Wang, Shichao Weng, Zhiqiang Wang, Xinyu Liu, Yang Yi, Peilong Zhou, Hongwei Tang
arXiv AI
3d ago

MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference

MaskCoFT introduces a masked co‑adaptive fine‑tuning approach for mixture‑of‑experts language models, training both routers and experts jointly with a cross‑entropy loss. A learnable binary mask limits each layer’s Top‑K routing to a subset of experts during fine‑tuning, allowing experts to adapt to the tokens they receive. In simulated GPU cache scenarios, MaskCoFT reduces expert fetches per token by 23.7% for Mixtral‑8x7B and 10.1% for DeepSeek‑V2‑Lite, and lowers inference time per output token by up to 16.4% and 5.5% respectively, while maintaining or improving accuracy across nine benchmarks.

By Junfeng Wu, Zehao Fan, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu, Yinan Wang
arXiv Machine Learning
Aug 27

Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing

The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.

By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng
arXiv Machine Learning
Sep 4

Towards a Statistical Understanding of Mixture-of-Experts

The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.

By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang