arXiv Machine Learning

Distributed Training using an Intelligent Network

The paper proposes using the network itself to aid distributed training over a wide area network (WAN). It suggests employing multicast for outbound traffic and in‑line FPGAs for inbound traffic to reduce bottlenecks, extending techniques commonly used in data centers to the WAN. An optimization framework generates synchronization schedules—rotating cliques of compute islands—tailored to the network topology, and demonstrates these ideas on a nine‑city WAN model, showing how schedules adapt to network capabilities.

arXiv Machine Learning
Aug 7

ML-for-ML

arXiv:2608. 06046v1 Announce Type: cross Abstract: AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important.

By Yutong Zhao, Noga H. Rotman, Gianni Antichi, Ran Ben Basat
arXiv Machine Learning
Sep 7

Tuning Collective Patterns to Alleviate Congestion in Shared AI Clusters

The paper introduces REACT, a system that dynamically tunes communication collectives in distributed AI training to mitigate congestion without requiring network infrastructure changes. REACT operates at the application layer, detecting congestion via flow statistics and adjusting the pattern of data exchange—such as selecting different aggregation nodes in an AllReduce tree—while preserving the semantics of the communication. Evaluations on a shared academic GPU cluster show that REACT improves algorithm bandwidth by 13%–38% under congestion, with simulations indicating potential gains up to 75%.

By Eashan Gupta, Yongzhou Chen, Apoorve Mohan, Pavlos Maniotis, Abdullah Kayi, Radhika Mittal
arXiv AI
Aug 26

ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork

ShardMeter is a lightweight analytical performance model that predicts end-to-end runtime for transformer-based workloads across sharded, distributed, and decentralized training setups. By taking a model’s characteristics and a target hardware topology as input, it estimates per-GPU and per-island throughput, training cost, total wall-clock time, and pinpoints performance bottlenecks. The model reveals diminishing-return regimes with increasing island size, quantifies compute- versus communication-bound scaling, evaluates hyperparameter trade-offs, and models cost-throughput for large-scale decentralized training, enabling rapid exploration of configuration space and near-optimal deployment plans.

By Tim Beringer (Technical University of Darmstadt), Patrick Diem (Technical University of Darmstadt), Felix Wolf (Technical University of Darmstadt), Arya Mazaheri (Technical University of Darmstadt, PanocularAI)
arXiv Machine Learning
Jul 16

Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models

arXiv:2607. 13332v1 Announce Type: new Abstract: Training large language models at the multi-billion to trillion parameter scale is confined to datacenters, where data-parallel (DP) and model-parallel (MP) techniques presume homogeneous accelerators, high-speed interconnects, and a single orchestrating entity.

By Gil Avraham, Violetta Shevchenko, Hadi Mohaghegh Dolatabadi, Karol Pajak, James Snewin, Harry Xi, Rodney O'Donnell, Thalaiyasingam Ajanthan, Sameera Ramasinghe, Chamin Hewa Koneputugodage, Shamane Siriwardhana, Alexander Long
arXiv AI
Sep 15

MANE: A Multi-Path Adaptive Network for Edge Onloading of Deep Neural Networks

MANE is a distributed inference framework that uses a multi‑path tail architecture to allow dynamic accuracy–throughput trade‑offs during edge onloading of deep neural networks. It introduces a novel multi‑path model, a three‑stage training scheme with Joint Head Network Distillation loss, and a hysteresis‑based scheduler with an equitable device‑fallback policy. The system achieves over 80% SLO satisfaction and 6pp higher accuracy than on‑device alternatives while supporting up to 40 concurrent devices.

By Sokratis Nikolaidis, Stylianos I. Venieris, Leonidas Malachias, Iakovos S. Venieris