arXiv AI By Qianli Liu, Kaibin Guo, Zicong Hong, Peng Li, Fahao Chen, Haodong Wang, Jian Lin, Song Guo

Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement

Read the original on arXiv AI →

arXiv:2607. 08782v1 Announce Type: cross Abstract: Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.