arXiv:2609.13585v1 Announce Type: cross
Abstract: Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granul...
By Ziming Mao, Yihan Zhang, Shawn Wei Chew, Shuang Ma, Costin Raiciu, Yang Zhou, Scott Shenker, Ion Stoica
arXiv:2606. 10493v1 Announce Type: cross Abstract: Local deployment of large Mixture-of-Experts (MoE) models falls short of the service quality achieved in cloud-scale environments, even under low-concurrency workloads.
By Wenxin Wang, Yule Hou, Yu Ji, Peng Qu, Youhui Zhang
arXiv:2609.38090v1 Announce Type: new
Abstract: Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-...
By Sanjali Yadav, Bahar Asgari
arXiv:2609.40093v1 Announce Type: cross
Abstract: Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substa...
By Jaehwan Lee, Sangmin Lee, Chaewon Kim, Junsik Shin, Jaejin Lee
arXiv:2606. 09200v1 Announce Type: cross Abstract: The rapid growth of large-scale machine learning (ML) has made distributed training across multiple GPUs a fundamental component of modern ML systems.
By Minyu Cui, Miquel Pericas
arXiv:2512. 10236v2 Announce Type: replace-cross Abstract: Modern ML workloads demand distributing training and inference across multiple GPUs.
By Shagnik Pal, Shaizeen Aga, Suchita Pati, Mahzabeen Islam, Lizy K. John