Communication-efficient pre-training of LLMs is increasingly important as training draws on compute distributed across clusters, data centers, and lower-bandwidth links. Many practical methods reduce communication frequency but still rely on synchronous All-Reduce operations that maintain identical model states and tie progress to global collectives.
arXiv:2508. 15706v3 Announce Type: replace Abstract: Communication-efficient distributed training algorithms (e.
By Amir Sarfi, Benjamin Th\'erien, Joel Lidin, Eugene Belilovsky
arXiv:2506.10911v2 Announce Type: replace
Abstract: Training large language models is generally done on clusters containing thousands of accelerators, communicating over a high-bandwidth interconnect...
By Jari Kolehmainen, Nikolay Blagoev, Semih Kara, John Donaghy, Christopher Nies, O\u{g}uzhan Ersoy
SeedFlood is a novel decentralized fine‑tuning method for large language models that scales to billions of parameters and hundreds of clients. It leverages the seed‑reconstructible structure of zeroth‑order gradients to reduce message sizes to near‑zero, enabling efficient flooding across the network. Experiments show SeedFlood outperforms standard zeroth‑order baselines in communication efficiency and generalization, and rivals first‑order gossip methods while incurring far less communication cost.
By Jihun Kim, Dongyeop Lee, Namhoon Lee
arXiv:2609.36662v1 Announce Type: cross
Abstract: The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of acc...
By Pengyu He, Yan Zhang, Ruien Li, Guangwen Yang
arXiv:2411. 02908v2 Announce Type: replace Abstract: Scaling large language models (LLMs) demands extensive data and computing resources, which are traditionally constrained to data centers by the high-bandwidth requirements of distributed training.
By Lorenzo Sani, Alex Iacob, Zeyu Cao, Royson Lee, Bill Marino, Yan Gao, Dongqi Cai, Zexi Li, Wanru Zhao, Xinchi Qiu, Nicholas D. Lane