arXiv:2608. 07226v1 Announce Type: cross Abstract: Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class accelerators remains limited.
By Vasanth Iyer
arXiv:2608. 05944v1 Announce Type: cross Abstract: We report operational experience full-fine-tuning a 32.
By Seon Ho Kim, Ui Jeong Jeon, Su Hyeon Kim, Min Tae Hwang
We report operational experience full-fine-tuning a 32. 76B-parameter dense model (Qwen3-32B) on 16 x NVIDIA B300 (two nodes, FSDP / ZeRO-3) -- among the first published field accounts on this accelerator.
arXiv:2607. 01409v1 Announce Type: cross Abstract: GPU training jobs fail often, roughly two in five on large production clusters, yet the operator typically learns of a failure only by reconnecting hours later.
By Parv Agarwal, Asif Ekbal
Leto is a fault‑tolerant training system for large language models that enables fast in‑place recovery on surviving hardware after hardware‑operable failures. It retains the working model state and reusable process state, while pre‑initializing remaining state in a shadow trainer, using two‑tier erasure protection and chunk‑level transactional updates to maintain consistency. Experiments on NVIDIA A100 clusters show Leto recovers 3.6–6.5× faster than checkpointing baselines and boosts productive training time by up to 13.7 percentage points, with simulations indicating over 95% productivity on a 131,072‑GPU cluster.
By Geon-Woo Kim, Joon Ha Kim, Daehyeok Kim
arXiv:2609.14762v1 Announce Type: cross
Abstract: Cloud-hosted large language models (LLMs) are increasingly used for root cause analysis (RCA) in AIOps pipelines, but they introduce data privacy ris...
By Rohit Patel, Susil Kumar Mohanty, Jeenal Chaudhary