Techniques for training large neural networks
Large neural networks are at the core of many recent advances in AI, but training them is a difficult engineering and research challenge which requires orchestrating a cluster of GPUs to perform a single synchronized calculation.
Related stories
HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models Training
arXiv:2609.39350v1 Announce Type: cross Abstract: As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training para...
HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models Training
As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training parallelism strategies at low cost while achieving supe...
Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems
The paper reports an empirical scalability study of data‑parallel training for Kolmogorov‑Arnold Networks (KANs) on high‑performance computing systems. Using up to eight NVIDIA A100 GPUs across four nodes on the FinisTerrae III supercomputer, the authors evaluate strong and weak scaling, communication overhead, and model‑size scaling, finding a 74.7% parallel efficiency and a 5.97× speedup at eight GPUs. They observe non‑monotonic communication costs driven by All‑Reduce choices and inter‑node latency, and note that while the parameter‑to‑memory ratio improves with larger models, training time scales less favorably, leading to guidelines for GPU topology and model‑size selection.
Mixture-of-Kittens: MoE Megakernel for NVL72s
arXiv:2609.36070v1 Announce Type: cross Abstract: AI accelerator systems are rapidly consolidating into scale-up architectures, where tens to thousands of GPUs communicate over high-bandwidth, single...
Enhancing AI Interpretability and Safety through Localised Architectures
arXiv:2606. 07998v1 Announce Type: cross Abstract: Recent advances in generative AI, especially powerful Large Language Models (LLMs) and Large Reasoning Models (LRMs), raise concerns over the interpretability, safety and sustainability of these large and opaque AI models.
Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models
arXiv:2607. 01844v1 Announce Type: cross Abstract: This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models.
Optimizing Energy-based Neural Network Training with Coherent Ising Machine
arXiv:2606. 09117v1 Announce Type: cross Abstract: While Ising machines serve as advanced physical solvers for the Ising model,enabling applications in combinatorial optimization and neural network training,their scalability for large-scale neural networks remains constrained by hardware connectivity limitations and suboptimal training methodologies.
Concurrent training methods for Kolmogorov-Arnold networks: Disjoint datasets and FPGA implementation
arXiv:2512. 18921v5 Announce Type: replace Abstract: The present paper introduces concurrency-driven enhancements to the training algorithm for the Kolmogorov-Arnold networks (KANs) that is based on the Newton-Kaczmarz (NK) method.
Enhancing AI Interpretability with Localised Architectures
arXiv:2606. 07998v3 Announce Type: replace-cross Abstract: Recent advances in generative AI, especially powerful Large Language Models (LLMs), raise concerns over the interpretability, safety and sustainability of these large and opaque AI models.
NoLoCo: No-all-reduce Low Communication Training Method for Large Models
arXiv:2506.10911v2 Announce Type: replace Abstract: Training large language models is generally done on clusters containing thousands of accelerators, communicating over a high-bandwidth interconnect...