arXiv:2402. 09589v2 Announce Type: replace-cross Abstract: We present MLCC, a novel technique to augment today's congestion control algorithms to accelerate DNN training jobs in shared GPU clusters in a fully distributed manner.
By Anton A. Zabreyko, Sanjoli Narang, Sudarsanan Rajasekaran, Manya Ghobadi
arXiv:2506. 01260v2 Announce Type: replace Abstract: Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to communication bottlenecks.
By Sameera Ramasinghe, Thalaiyasingam Ajanthan, Gil Avraham, Yan Zuo, Alexander Long
arXiv:2508. 15706v3 Announce Type: replace Abstract: Communication-efficient distributed training algorithms (e.
By Amir Sarfi, Benjamin Th\'erien, Joel Lidin, Eugene Belilovsky
arXiv:2410. 05662v4 Announce Type: replace Abstract: Most federated learning (FL) approaches assume a fixed device set.
By Zhan-Lun Chang, Dong-Jun Han, Seyyedali Hosseinalipour, Mung Chiang, Christopher G. Brinton
arXiv:2608. 06046v1 Announce Type: cross Abstract: AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important.
By Yutong Zhao, Noga H. Rotman, Gianni Antichi, Ran Ben Basat
arXiv:2607. 24692v1 Announce Type: cross Abstract: Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline together with a higher-accuracy slow path that runs higher-compute methods on stronger, remote hardware, so its results can be returned on time and combined with the fast path predictions.
By Jhonatan Tavori, Gur-Eyal Sela, Ion Stoica, Gil Zussman