arXiv:2605. 09370v3 Announce Type: replace-cross Abstract: Large-scale AI training is now fundamentally a distributed systems problem, and hardware failures have become routine operating conditions rather than rare exceptions.
By Daemyung Kang, Eunjin Hwang, Hanjeong Lee, HyeokJin Kim, Hyunhoi Koo, Jeongkyu Shin, Jeongseok Kang, Jihyun Kang, Joongi Kim, Junbum Lee, Jungseung Yang, Kyujin Cho, Youngsook Song
OpenAI and NVIDIA announce a strategic partnership to deploy 10 gigawatts of AI datacenters powered by NVIDIA systems, with the first phase launching in 2026.
Leto is a fault‑tolerant training system for large language models that enables fast in‑place recovery on surviving hardware after hardware‑operable failures. It retains the working model state and reusable process state, while pre‑initializing remaining state in a shadow trainer, using two‑tier erasure protection and chunk‑level transactional updates to maintain consistency. Experiments on NVIDIA A100 clusters show Leto recovers 3.6–6.5× faster than checkpointing baselines and boosts productive training time by up to 13.7 percentage points, with simulations indicating over 95% productivity on a 131,072‑GPU cluster.
By Geon-Woo Kim, Joon Ha Kim, Daehyeok Kim
A measured look at distributed training, from DDP and FSDP to the ZeRO stages in between, and why the wiring between your GPUs matters as much as the strategy you choose The post Behind the Scenes of Distributed Training and Why Your GPU Wiring Matters as Much as Your Strategy appeared first on Towards Data Science .
By Hussen Mohammed Ibrahim
The paper introduces REACT, a system that dynamically tunes communication collectives in distributed AI training to mitigate congestion without requiring network infrastructure changes. REACT operates at the application layer, detecting congestion via flow statistics and adjusting the pattern of data exchange—such as selecting different aggregation nodes in an AllReduce tree—while preserving the semantics of the communication. Evaluations on a shared academic GPU cluster show that REACT improves algorithm bandwidth by 13%–38% under congestion, with simulations indicating potential gains up to 75%.
By Eashan Gupta, Yongzhou Chen, Apoorve Mohan, Pavlos Maniotis, Abdullah Kayi, Radhika Mittal
arXiv:2608. 07226v1 Announce Type: cross Abstract: Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class accelerators remains limited.
By Vasanth Iyer
The paper introduces Global Clustered Parallel Split Learning (GCPSL), a method that groups clients into fixed clusters and runs Parallel Split Learning with Global Sampling (GPSL) concurrently across these clusters, periodically merging client and server model segments. Experiments with 256 logical clients show that distributing the population across more workloads increases direct data participation, though smaller clusters may slightly reduce accuracy. In a practical four‑GPU setup, label‑aware GCPSL achieves 85 % CIFAR‑10 validation accuracy in roughly 6 minutes, compared to over 19 minutes for serialized workloads, with size‑balanced cluster assignments improving participation by 3.25 percentage points.
By Mohammad Kohankhaki, Valentin Rentschler, Anke Schmeink