arXiv Machine Learning By Nihar Shah, Ben Blier

Distributed Training using an Intelligent Network

Read the original on arXiv Machine Learning →

The paper proposes using the network itself to aid distributed training over a wide area network (WAN). It suggests employing multicast for outbound traffic and in‑line FPGAs for inbound traffic to reduce bottlenecks, extending techniques commonly used in data centers to the WAN. An optimization framework generates synchronization schedules—rotating cliques of compute islands—tailored to the network topology, and demonstrates these ideas on a nine‑city WAN model, showing how schedules adapt to network capabilities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 7

ML-for-ML

arXiv:2608. 06046v1 Announce Type: cross Abstract: AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important.

By Yutong Zhao, Noga H. Rotman, Gianni Antichi, Ran Ben Basat
arXiv Machine Learning
Sep 7

Tuning Collective Patterns to Alleviate Congestion in Shared AI Clusters

The paper introduces REACT, a system that dynamically tunes communication collectives in distributed AI training to mitigate congestion without requiring network infrastructure changes. REACT operates at the application layer, detecting congestion via flow statistics and adjusting the pattern of data exchange—such as selecting different aggregation nodes in an AllReduce tree—while preserving the semantics of the communication. Evaluations on a shared academic GPU cluster show that REACT improves algorithm bandwidth by 13%–38% under congestion, with simulations indicating potential gains up to 75%.

By Eashan Gupta, Yongzhou Chen, Apoorve Mohan, Pavlos Maniotis, Abdullah Kayi, Radhika Mittal