Towards Data Science

Behind the Scenes of Distributed Training and Why Your GPU Wiring Matters as Much as Your Strategy

A measured look at distributed training, from DDP and FSDP to the ZeRO stages in between, and why the wiring between your GPUs matters as much as the strategy you choose The post Behind the Scenes of Distributed Training and Why Your GPU Wiring Matters as Much as Your Strategy appeared first on Towards Data Science .

arXiv Machine Learning
Sep 15

Communication-Efficient LLM Adaptation over Decentralized GPU Meshes

The paper introduces a communication‑efficient method for adapting large language models on decentralized GPU meshes. It proposes an asynchronous two‑circuit system that uses fast compressed training with activation masking for pipeline‑parallel transfer and compressed data‑parallel synchronization, while a slower anchor circuit performs occasional unmasked passes. A spectral correction optimizer then denoises the masked gradients using these anchor priors, enabling high compression rates and achieving up to 40× throughput gains over internet‑grade connections while matching dense uncompressed performance.

By Sameera Ramasinghe, Shamane Siriwardhana, Thalaiyasingam Ajanthan, Hadi Mohaghegh Dolatabadi, Chamin P Hewa Koneputugodage, Gil Avraham, Violetta Shevchenko, James Snewin, Karol Pajak, Harry Xi, Alexander Long
arXiv Machine Learning
Sep 10

Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems

The paper reports an empirical scalability study of data‑parallel training for Kolmogorov‑Arnold Networks (KANs) on high‑performance computing systems. Using up to eight NVIDIA A100 GPUs across four nodes on the FinisTerrae III supercomputer, the authors evaluate strong and weak scaling, communication overhead, and model‑size scaling, finding a 74.7% parallel efficiency and a 5.97× speedup at eight GPUs. They observe non‑monotonic communication costs driven by All‑Reduce choices and inter‑node latency, and note that while the parameter‑to‑memory ratio improves with larger models, training time scales less favorably, leading to guidelines for GPU topology and model‑size selection.

By Guangneng Chen, David Garcia Selfa, Pablo Quesada Barriuso
arXiv AI
Jun 9

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs

arXiv:2605. 09370v3 Announce Type: replace-cross Abstract: Large-scale AI training is now fundamentally a distributed systems problem, and hardware failures have become routine operating conditions rather than rare exceptions.

By Daemyung Kang, Eunjin Hwang, Hanjeong Lee, HyeokJin Kim, Hyunhoi Koo, Jeongkyu Shin, Jeongseok Kang, Jihyun Kang, Joongi Kim, Junbum Lee, Jungseung Yang, Kyujin Cho, Youngsook Song
arXiv AI
Jul 7

Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

arXiv:2602. 24044v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of adapters must be hosted concurrently.

By Ferran Agullo, Joan Oliveras, Chen Wang, Alberto Gutierrez-Torre, Olivier Tardieu, Alaa Youssef, Jordi Torres, Josep Ll. Berral
Hugging Face Trending Papers
Sep 17

Accelerating Sharded Data Parallelism at Scale with Federated Learning

The paper proposes two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning-style aggregations to reduce communication overhead in large-scale AI training. By forming loosely‑coupled federation groups, the methods keep inter‑group traffic minimal while maintaining a bounded global batch size. Experiments on a Llama3.1 8B model trained on 512 A100 GPUs show up to 8.04× faster data processing and 4.48 lower evaluation perplexity compared to traditional sharded DP approaches.