OpenAI Blog

Techniques for training large neural networks

Read the original on OpenAI Blog →

Large neural networks are at the core of many recent advances in AI, but training them is a difficult engineering and research challenge which requires orchestrating a cluster of GPUs to perform a single synchronized calculation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at OpenAI Blog.

arXiv Machine Learning
Sep 10

Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems

The paper reports an empirical scalability study of data‑parallel training for Kolmogorov‑Arnold Networks (KANs) on high‑performance computing systems. Using up to eight NVIDIA A100 GPUs across four nodes on the FinisTerrae III supercomputer, the authors evaluate strong and weak scaling, communication overhead, and model‑size scaling, finding a 74.7% parallel efficiency and a 5.97× speedup at eight GPUs. They observe non‑monotonic communication costs driven by All‑Reduce choices and inter‑node latency, and note that while the parameter‑to‑memory ratio improves with larger models, training time scales less favorably, leading to guidelines for GPU topology and model‑size selection.

By Guangneng Chen, David Garcia Selfa, Pablo Quesada Barriuso
arXiv Machine Learning
4d ago

Mixture-of-Kittens: MoE Megakernel for NVL72s

arXiv:2609.36070v1 Announce Type: cross Abstract: AI accelerator systems are rapidly consolidating into scale-up architectures, where tens to thousands of GPUs communicate over high-bandwidth, single...

By Stuart H. Sul, Nash Brown, Henry Wildermuth, William Lin, Federico Cassano, Christopher R\'e