arXiv Machine Learning By Vu Quang Hoang, Nghia Hieu Nguyen

Dense Mixture-of-Experts as a Reparameterized Wide FFN: A Granularity Sweep at Fixed Compute

Read the original on arXiv Machine Learning →

The paper investigates a dense analogue of Sparse Mixture-of-Experts (MoE) models by using $K$ SwiGLU experts that are all active for every token and combined via a softmax gate, keeping the total feed‑forward network (FFN) width fixed. Validation loss shows a non‑monotonic relationship with $K$: $K=2$ slightly improves performance over the single‑expert baseline, while $K=4$ and $K=6$ degrade it. The study also finds that the gating mechanism remains largely soft and balanced, except for a near one‑hot routing in the first layer of the $K=4$ model, which is functionally important as forcing uniform routing increases loss significantly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 28

Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs

MetaNet is a support‑set controller that predicts, for each layer of a Mixture‑of‑Experts model, an expert‑retention threshold and a bounded routing bias while keeping the backbone, experts, and router frozen. On DeepSeek‑MoE‑16B‑Chat, MetaNet offers a tunable trade‑off between accuracy and expert activation: a conservative setting activates 3.61 experts on average (40% fewer than a fixed k=6) with comparable MMLU accuracy, whereas an aggressive setting activates only 2.28 experts (62% fewer) with a modest accuracy drop. The MMLU‑trained controller also transfers to C‑Eval, activating 2.90 experts on average (52% fewer than fixed k=6) at 0.386 accuracy.

By Rongfeng Wang, Shichao Weng, Zhiqiang Wang, Xinyu Liu, Yang Yi, Peilong Zhou, Hongwei Tang
arXiv Machine Learning
Aug 27

Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing

The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.

By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng