The paper introduces Federation of Experts (FoE), a new architecture that reorganizes the mixture-of-experts (MoE) block in transformer layers into multiple MoE clusters. Each cluster handles a single KV head, and expert parallelism is applied within clusters while a sum operation synchronizes post‑attention residuals across clusters. FoE eliminates all‑to‑all communication on a single GPU and limits it to intra‑node communication in multi‑node setups, leading to significant reductions in inference latency and throughput improvements on LongBench.
By Muhammad Shahir Abdurrahman, Chun Deng, Azalia Mirhoseini, Philip Levis
arXiv:2606. 10493v1 Announce Type: cross Abstract: Local deployment of large Mixture-of-Experts (MoE) models falls short of the service quality achieved in cloud-scale environments, even under low-concurrency workloads.
By Wenxin Wang, Yule Hou, Yu Ji, Peng Qu, Youhui Zhang
arXiv:2506.10911v2 Announce Type: replace
Abstract: Training large language models is generally done on clusters containing thousands of accelerators, communicating over a high-bandwidth interconnect...
By Jari Kolehmainen, Nikolay Blagoev, Semih Kara, John Donaghy, Christopher Nies, O\u{g}uzhan Ersoy
arXiv:2607. 19539v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language models (LLMs) to trillion-parameter regimes.
By Minyu Cui, Anna Wingkvist, Morgan Ericsson
arXiv:2609.36070v1 Announce Type: cross
Abstract: AI accelerator systems are rapidly consolidating into scale-up architectures, where tens to thousands of GPUs communicate over high-bandwidth, single...
By Stuart H. Sul, Nash Brown, Henry Wildermuth, William Lin, Federico Cassano, Christopher R\'e
arXiv:2603. 28768v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) has recently emerged as the mainstream architecture for efficiently scaling large language models while maintaining near-constant computational cost.
By Adrian Zhao, Zhenkun Cai, Zhenyu Song, Lingfan Yu, Haozheng Fan, Jun Wu, Yida Wang, Nandita Vijaykumar