Scaling Kubernetes to 7,500 nodes
We’ve scaled Kubernetes clusters to 7,500 nodes, producing a scalable infrastructure for large models like GPT-3, CLIP, and DALL·E, but also for rapid small-scale iterative research such as Scaling Laws for Neural Language Models.
Related stories
Scaling laws for neural language models
Kolmogorov--Arnold Networks for Small Language Models
arXiv:2607. 15525v1 Announce Type: cross Abstract: Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks.
Techniques for training large neural networks
Large neural networks are at the core of many recent advances in AI, but training them is a difficult engineering and research challenge which requires orchestrating a cluster of GPUs to perform a single synchronized calculation.
Towards Engineering Scaling Laws with Pretraining Data Composition
arXiv:2606. 19781v1 Announce Type: cross Abstract: Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size.
Scaling Laws of Global Weather Models
arXiv:2602. 22962v2 Announce Type: replace Abstract: Data-driven models are revolutionizing weather forecasting.
Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems
The paper reports an empirical scalability study of data‑parallel training for Kolmogorov‑Arnold Networks (KANs) on high‑performance computing systems. Using up to eight NVIDIA A100 GPUs across four nodes on the FinisTerrae III supercomputer, the authors evaluate strong and weak scaling, communication overhead, and model‑size scaling, finding a 74.7% parallel efficiency and a 5.97× speedup at eight GPUs. They observe non‑monotonic communication costs driven by All‑Reduce choices and inter‑node latency, and note that while the parameter‑to‑memory ratio improves with larger models, training time scales less favorably, leading to guidelines for GPU topology and model‑size selection.
Federation of Experts: Communication Efficient Distributed Inference for Large Language Models
The paper introduces Federation of Experts (FoE), a new architecture that reorganizes the mixture-of-experts (MoE) block in transformer layers into multiple MoE clusters. Each cluster handles a single KV head, and expert parallelism is applied within clusters while a sum operation synchronizes post‑attention residuals across clusters. FoE eliminates all‑to‑all communication on a single GPU and limits it to intra‑node communication in multi‑node setups, leading to significant reductions in inference latency and throughput improvements on LongBench.
Branch Scaling Manifests as Implicit Architectural Regularization for Improving Generalization in Overparameterized ResNets
arXiv:2403.04545v4 Announce Type: replace Abstract: Scaling factors in residual branches have emerged as a prevalent method for boosting neural network performance, especially in normalization-free a...
Kilobyte Models: Neural Networks as a Seed and a Quantized Latent
arXiv:2608. 00860v1 Announce Type: new Abstract: The cost of storing and transmitting a trained neural network scales with its parameter count, a bottleneck for over-the-air updates, on-device libraries, and other bandwidth-bound deployments.
How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size
arXiv:2607. 01487v1 Announce Type: new Abstract: We propose a scaling law that takes into account model size and training data while explicitly splitting the latter into training steps and batch size (called three-term law).
FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs
arXiv:2606. 19025v1 Announce Type: cross Abstract: Pre-training Large Language Models (LLMs) typically demands large-scale infrastructure with tightly coupled hardware accelerators.