arXiv Machine Learning

MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs

arXiv:2510. 19366v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) scales model capacity through sparse activation, and is becoming an important architecture for large language models (LLMs).