Hugging Face Blog

Introducing Training Cluster as a Service - a new collaboration with NVIDIA

arXiv AI
Jun 9

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs

arXiv:2605. 09370v3 Announce Type: replace-cross Abstract: Large-scale AI training is now fundamentally a distributed systems problem, and hardware failures have become routine operating conditions rather than rare exceptions.

By Daemyung Kang, Eunjin Hwang, Hanjeong Lee, HyeokJin Kim, Hyunhoi Koo, Jeongkyu Shin, Jeongseok Kang, Jihyun Kang, Joongi Kim, Junbum Lee, Jungseung Yang, Kyujin Cho, Youngsook Song
arXiv Machine Learning
3d ago

Leto: Fast In-Place Recovery for LLM Training on Surviving Hardware

Leto is a fault‑tolerant training system for large language models that enables fast in‑place recovery on surviving hardware after hardware‑operable failures. It retains the working model state and reusable process state, while pre‑initializing remaining state in a shadow trainer, using two‑tier erasure protection and chunk‑level transactional updates to maintain consistency. Experiments on NVIDIA A100 clusters show Leto recovers 3.6–6.5× faster than checkpointing baselines and boosts productive training time by up to 13.7 percentage points, with simulations indicating over 95% productivity on a 131,072‑GPU cluster.

By Geon-Woo Kim, Joon Ha Kim, Daehyeok Kim
arXiv Machine Learning
Sep 7

Tuning Collective Patterns to Alleviate Congestion in Shared AI Clusters

The paper introduces REACT, a system that dynamically tunes communication collectives in distributed AI training to mitigate congestion without requiring network infrastructure changes. REACT operates at the application layer, detecting congestion via flow statistics and adjusting the pattern of data exchange—such as selecting different aggregation nodes in an AllReduce tree—while preserving the semantics of the communication. Evaluations on a shared academic GPU cluster show that REACT improves algorithm bandwidth by 13%–38% under congestion, with simulations indicating potential gains up to 75%.

By Eashan Gupta, Yongzhou Chen, Apoorve Mohan, Pavlos Maniotis, Abdullah Kayi, Radhika Mittal
arXiv Machine Learning
Sep 25

Concurrent Split Learning Through Stable Client Clustering

The paper introduces Global Clustered Parallel Split Learning (GCPSL), a method that groups clients into fixed clusters and runs Parallel Split Learning with Global Sampling (GPSL) concurrently across these clusters, periodically merging client and server model segments. Experiments with 256 logical clients show that distributing the population across more workloads increases direct data participation, though smaller clusters may slightly reduce accuracy. In a practical four‑GPU setup, label‑aware GCPSL achieves 85 % CIFAR‑10 validation accuracy in roughly 6 minutes, compared to over 19 minutes for serialized workloads, with size‑balanced cluster assignments improving participation by 3.25 percentage points.

By Mohammad Kohankhaki, Valentin Rentschler, Anke Schmeink