arXiv:2402. 09589v2 Announce Type: replace-cross Abstract: We present MLCC, a novel technique to augment today's congestion control algorithms to accelerate DNN training jobs in shared GPU clusters in a fully distributed manner.
By Anton A. Zabreyko, Sanjoli Narang, Sudarsanan Rajasekaran, Manya Ghobadi
arXiv:2606. 11081v1 Announce Type: cross Abstract: Communication-efficient pre-training of LLMs is increasingly important as training draws on compute distributed across clusters, data centers, and lower-bandwidth links.
By Pietro Cagnasso, Eugene Belilovsky, Edouard Oyallon
arXiv:2608. 06046v1 Announce Type: cross Abstract: AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important.
By Yutong Zhao, Noga H. Rotman, Gianni Antichi, Ran Ben Basat
The paper introduces REACT, a system that dynamically tunes communication collectives in distributed AI training to mitigate congestion without requiring network infrastructure changes. REACT operates at the application layer, detecting congestion via flow statistics and adjusting the pattern of data exchange—such as selecting different aggregation nodes in an AllReduce tree—while preserving the semantics of the communication. Evaluations on a shared academic GPU cluster show that REACT improves algorithm bandwidth by 13%–38% under congestion, with simulations indicating potential gains up to 75%.
By Eashan Gupta, Yongzhou Chen, Apoorve Mohan, Pavlos Maniotis, Abdullah Kayi, Radhika Mittal
Communication-efficient pre-training of LLMs is increasingly important as training draws on compute distributed across clusters, data centers, and lower-bandwidth links. Many practical methods reduce communication frequency but still rely on synchronous All-Reduce operations that maintain identical model states and tie progress to global collectives.
arXiv:2608. 05340v1 Announce Type: cross Abstract: Time-Sensitive Networking (TSN) and Mobile Edge Computing (MEC) hold strong potential for enabling ultra-reliable low-latency communication for time-sensitive applications, such as eXtended Reality (XR).
By Marcos Carvalho, Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci, Daniel F. Macedo
ShardMeter is a lightweight analytical performance model that predicts end-to-end runtime for transformer-based workloads across sharded, distributed, and decentralized training setups. By taking a model’s characteristics and a target hardware topology as input, it estimates per-GPU and per-island throughput, training cost, total wall-clock time, and pinpoints performance bottlenecks. The model reveals diminishing-return regimes with increasing island size, quantifies compute- versus communication-bound scaling, evaluates hyperparameter trade-offs, and models cost-throughput for large-scale decentralized training, enabling rapid exploration of configuration space and near-optimal deployment plans.
By Tim Beringer (Technical University of Darmstadt), Patrick Diem (Technical University of Darmstadt), Felix Wolf (Technical University of Darmstadt), Arya Mazaheri (Technical University of Darmstadt, PanocularAI)
arXiv:2607. 13332v1 Announce Type: new Abstract: Training large language models at the multi-billion to trillion parameter scale is confined to datacenters, where data-parallel (DP) and model-parallel (MP) techniques presume homogeneous accelerators, high-speed interconnects, and a single orchestrating entity.
By Gil Avraham, Violetta Shevchenko, Hadi Mohaghegh Dolatabadi, Karol Pajak, James Snewin, Harry Xi, Rodney O'Donnell, Thalaiyasingam Ajanthan, Sameera Ramasinghe, Chamin Hewa Koneputugodage, Shamane Siriwardhana, Alexander Long
arXiv:2608. 07733v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems.
By Liad Gerstman, Aditya Dhakal, Dejan Milojicic, Avi Mendelson
arXiv:2606. 19025v1 Announce Type: cross Abstract: Pre-training Large Language Models (LLMs) typically demands large-scale infrastructure with tightly coupled hardware accelerators.
By Lorenzo Sani, Zeyu Cao, Meghdad Kurmanji, Alex Iacob, Andrej Jovanovic, Yan Gao, Wanru Zhao, Nicholas D. Lane
MANE is a distributed inference framework that uses a multi‑path tail architecture to allow dynamic accuracy–throughput trade‑offs during edge onloading of deep neural networks. It introduces a novel multi‑path model, a three‑stage training scheme with Joint Head Network Distillation loss, and a hysteresis‑based scheduler with an equitable device‑fallback policy. The system achieves over 80% SLO satisfaction and 6pp higher accuracy than on‑device alternatives while supporting up to 40 concurrent devices.
By Sokratis Nikolaidis, Stylianos I. Venieris, Leonidas Malachias, Iakovos S. Venieris
arXiv:2605. 01989v2 Announce Type: replace Abstract: Distributed machine learning (ML) training has become a necessity with the prevalence of billion to trillion-parameter-scale models.
By Zechen Ma, Zixi Qu, Jinyan Yi, David Lin, Yashar Ganjali