arXiv Machine Learning By Geon-Woo Kim, Joon Ha Kim, Daehyeok Kim

Leto: Fast In-Place Recovery for LLM Training on Surviving Hardware

Read the original on arXiv Machine Learning →

Leto is a fault‑tolerant training system for large language models that enables fast in‑place recovery on surviving hardware after hardware‑operable failures. It retains the working model state and reusable process state, while pre‑initializing remaining state in a shadow trainer, using two‑tier erasure protection and chunk‑level transactional updates to maintain consistency. Experiments on NVIDIA A100 clusters show Leto recovers 3.6–6.5× faster than checkpointing baselines and boosts productive training time by up to 13.7 percentage points, with simulations indicating over 95% productivity on a 131,072‑GPU cluster.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 2

DeadPool: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint

State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware stack. Existing fault-tolerance mechanisms either impose non-trivial overhead during failure-free execution or suffer from prolonged recovery latency, particularly under scenarios where a small subset of compute nodes experience permanent failures.

arXiv AI
Jun 9

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs

arXiv:2605. 09370v3 Announce Type: replace-cross Abstract: Large-scale AI training is now fundamentally a distributed systems problem, and hardware failures have become routine operating conditions rather than rare exceptions.

By Daemyung Kang, Eunjin Hwang, Hanjeong Lee, HyeokJin Kim, Hyunhoi Koo, Jeongkyu Shin, Jeongseok Kang, Jihyun Kang, Joongi Kim, Junbum Lee, Jungseung Yang, Kyujin Cho, Youngsook Song
arXiv AI
Aug 5

Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

arXiv:2608. 03796v1 Announce Type: cross Abstract: Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD).

By Bakbergen Ryskulov, Iker Garc\'ia-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Rom\'an Or\'us