arXiv AI

DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts

arXiv:2606. 01062v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have become a leading approach for decoupling parameter count from computational cost in large language models, yet effectively scaling MoE performance remains a challenge.

arXiv Computation and Language
Sep 2

PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition

PCoMoE introduces a path‑compositional execution framework that moves Mixture‑of‑Experts inference from coarse‑grained expert selection to fine‑grained path composition. It uses a path‑level formulation of expert computation, a compatibility‑aware layer‑wise pruning strategy to eliminate low‑value path combinations, and a hardware‑friendly execution engine that reuses sub‑expert structures with bounded overhead. Experiments show up to a 1.31× speedup and a 10% accuracy improvement over existing MoE inference methods.

By Ziyan Gan, Fangxin Liu, Chenyang Guan, Junjie Wang, Ning Yang, Haomin Li, Xiang Li, Siran Yang, Jiamang Wang, Lin Qu, Zongwu Wang, Li Jiang, Haibing Guan
arXiv AI
Sep 17

MoRE: Mixture of Reused Experts

MoRE: Mixture of Reused Experts is a hybrid architecture that combines Mixture-of-Experts (MoE) with weight‑sharing techniques. It shares expert pools across adjacent layers while each layer keeps its own router, and introduces lightweight depth embeddings to help shared experts differentiate layer contexts. Experiments on models ranging from 114 M to 1.15 B parameters show MoRE achieves lower perplexity and better downstream performance than standard MoEs and other weight‑sharing models, with only minimal changes to existing MoE implementations.

By Eric S. Qiu, Utku Umur Acikalin, Justin Lovelace, Christian Belardi, Arjun B. Mulchandani, Carla P. Gomes, Kilian Q. Weinberger
arXiv Machine Learning
Aug 27

Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing

The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.

By Yuanteng Chen, Peisong Wang, Yuantian Shao, Nanxin Zeng, Chang Xu, Jian Cheng